Tier 03 · Intermediate

The Nexus

For the person about to download one, who wants to know which, and what it will need.

Every family in this tier gets the same treatment: what it does, how it works in a few sentences, the words on its cards, the models that anchor it on the day this guide was checked, what runs it, the memory it needs, the licence trap, and the one thing people get wrong. The anchors are a snapshot, and the date matters more here than anywhere else in the guide.

3.1

Language models: which variant are you looking at?

Gemma 4 ships in five sizes, and each size ships as a base model and an instruction-tuned one; Qwen3.8-27B ships with thinking on by default and a switch to turn it off; gpt-oss OpenAI ships with a reasoning-effort dial marked low, medium and high. Before you have chosen a family you are already choosing among variants.

You have met this: the picker, again. A chat product's picker is this choice with the filenames taken off: often a fast model and a thinking one, sometimes a small one and a large one. The card's words are the picker's words, and the rest of this section is what they mean. A coding agent may expose the same sort of choice; a coder variant, where a family ships one, is a separate checkpoint and not a setting.

How it works. The loop the Origin tier gave you: predict the next token from all the tokens so far, pick one, append it, ask again, until the model emits its stop token. A chat is that loop, wrapped in the template from the Vector tier.

The words. Two paragraphs here, the variants and the context; the next section takes what the model emits, the dials and the two speeds.

The variants. A base model is the raw result of pre-training; it continues text and is not built to answer. An instruct or chat model has been post-trained to answer; the -it on Gemma's Google DeepMind names and -Instruct on Llama's Meta mean this, and Qwen's plain Qwen3.8-27B is already post-trained according to its card. A thinking or reasoning model spends tokens on working before its answer, which may be shown, hidden or summarised depending on the system; Qwen's Alibaba card calls it thinking mode and makes it default, gpt-oss calls it reasoning effort, and Gemma 4 offers a token budget for it in five fixed steps from 70 to 1,120 tokens. That working is output tokens, as the Vector tier said, and gpt-oss's card is explicit that it is not meant to be shown to end users. Vision or VL variants take images; on the hub the task tag image-text-to-text marks them, and Qwen3.8-27B, Gemma 4 and both Llama 4 models carry it. Coder variants exist as separate checkpoints in most families.

Context. The window is a token count from the card: 262,144 for Qwen3.8-27B and the larger Gemma 4 models, 131,072 for the Gemma 4 edge models and gpt-oss, 1,048,576 for DeepSeek V4 and Kimi K3 Moonshot; GLM-5.3's config exposes the same figure and its card evaluates at 1M without stating a headline context length. It is a limit on what the model can see, not on what it knows, and a bigger window costs memory for the cache that holds it. That cache, the KV cache, has its own guide; here it is enough to know it exists and grows with context. It is also request state: each active request needs cache space of its own, apart from any prefix the runtime can share, so the weights stay fixed while the cache grows with every request being served at once, and a server can become limited by cache space even though the weights are shared. vLLM prints both when it starts, the cache's capacity in tokens and an estimate of how many requests of a given length it can serve at once.

The anchors. Licence and gating as read from the hub on the day this guide was checked. Total parameters as counted in the safetensors files (the file), with the card's own total where it differs; active parameters and routing from the card; context as the maker advertises it on the card, which is the supported figure, since a config can expose a larger number than the maker will stand behind.

Model Total (file) Active (card) Experts and routing Context (card) Licence Gated
Qwen3.8-27B 27.8B dense none 262,144 native, 1M extended Apache 2.0 no
Qwen3.8-Flash-Next 180B (125B model, 51B n-gram embedding, 4B MTP) 6B 512 experts; 10 routed + 1 shared 262,144 "other", qwen-community-1.0 no
DeepSeek-V4.1-Flash 763B (552B backbone, 196B Engram memory) 8B / 16B 384 experts; 6 per token 1,048,576 MIT no
DeepSeek-V4-Pro-0813 1,651B (1.6T backbone) 49B per the V4.1 card's table 384 experts; 6 per token 1,048,576 MIT no
Kimi-K3 2,780B (2.8T) 104B 896 experts; 16 selected + 2 shared 1,048,576 "other", kimi-k3 no
GLM-5.3 753B not stated on card 256 experts; 8 per token 1M used in card evaluations; 1,048,576 in config "other", glm-5.3 no
Llama 4 Scout 109B 17B 16 experts 10M "other", llama4 manual
Llama 4 Maverick 402B (400B) 17B 128 experts 1M "other", llama4 manual
Gemma 4 26B-A4B 25.8B (25.2B) 3.8B 128 experts; 8 active + 1 shared 256K Apache 2.0 no
Gemma 4 12B 12.0B dense none 256K Apache 2.0 no
Gemma 4 E4B / E2B 8.0B / 5.1B 4.5B / 2.3B "effective" none 128K Apache 2.0 no
gpt-oss-120b / 20b 117B / 21B 5.1B / 3.6B MoE, per card 131,072 Apache 2.0 no
Mistral Small 4 119B 6.5B (branded A6B) 128 experts; 4 active 256K Apache 2.0 no
Phi-4 14.7B dense none 16K MIT no
Nemotron 3 Nano 30B-A3B 31.6B 3B per its name 128 experts; 6 per token 262,144 "other", nvidia-nemotron-open-model-license no

Three things in that table to notice. Almost every large model in this table is a mixture of experts, and the active counts are small next to the totals: Kimi's 2.8 trillion runs 104 billion per token, 16 of 896 experts plus 2 shared. The million-token context sits at the top of the table and Llama 4 Scout claims ten million. And the Gemma edge models carry two parameter counts, because the "E" means effective: the card says the E2B file holds 5.1 billion parameters of which 2.3 billion are counted as effective, the rest being per-layer embeddings kept out of the count. The trade's word for the small end of the table is small language model, or SLM, and it is a size, not a kind.

What runs it. The Qwen card names Transformers, vLLM, SGLang and TokenSpeed, and says the last three are for production and high throughput. The GGUF world, llama.cpp and the tools built on it (Ollama and LM Studio among them), is optimised first for local inference on one machine, though it can serve an API to several users; vLLM and SGLang are designed around high-throughput serving and batching. Both worlds advertise an OpenAI-compatible endpoint, an address that accepts requests in the shape one provider's API made common, which means a basic chat call can move between them, or to a provider's API, with little change; tool calling, reasoning controls, image input and structured-output extensions (4.6) differ between providers, and those parts of your code will not move as cleanly.

Memory floor. The weights, plus the KV cache for the context you will use, plus the runtime's own working space; a 15.5 GB file does not comfortably fit a 16 GB card. For Qwen3.8-27B: 16.46 GB for the Q4_K_M plus 0.93 GB for the vision projector, before any context. gpt-oss's own card gives its floors directly: the 120b on a single 80 GB card, the 20b within 16 GB, both because the MoE weights were post-trained at MXFP4. MXFP4 is the four-bit floating-point member of the Microscaling (MX) formats, which pair a narrow number type with a scale shared across a block of values; gpt-oss ships its MoE weights in it from the maker, rather than as a later third-party quantisation. For everything else, read the file sizes in the quantiser's folder and add room for the window you intend to use. Where the quantiser publishes quality measurements, look for perplexity or KL divergence against the BF16 or base model; that comparison is the measured loss the Vector tier meant. When the weights do not fit on one accelerator, runtimes can split the model across several; tensor parallelism, named at 2.12, is one common way to do it.

The licence trap. The same family carries permissive and restricted checkpoints side by side. On the day, Qwen3.8-27B was Apache 2.0 and Qwen3.8-Flash-Next, released two weeks later by the same lab, was under a Qwen community licence. "Qwen is Apache" is not a sentence; "this checkpoint is Apache" is.

What people get wrong. Choosing by leaderboard rank. The rank measures the vendor's benchmarks on the vendor's settings, and a smaller model that fits your card with room for context beats a larger one that spills. Choose by fit, then by the languages you need (Gemma 4's card claims 35+ supported and 140+ in pre-training; test yours), then by the task, and only then by the table.

New here: attention, the KV cache and the Transformer. How the next-token prediction is computed, why context costs memory, and what "Gated DeltaNet" on the Qwen card means belong to neural networks, and to attention and the KV cache. Nothing here needs them.

3.2

Language models: the dials and the two speeds

What it emits. Before the dials make sense, one paragraph on what the model produces. For each position it emits a score, a logit, for every token in its vocabulary. Those scores are turned into a probability distribution. Then a token is chosen from it: greedy decoding takes the most probable one every time and is deterministic; a sampler draws from the distribution, and temperature, top-p, min_p and the penalties reshape the distribution before the draw. That draw is why the same prompt gives a different answer twice; a fixed seed improves reproducibility on the same runtime and hardware, but kernels, batching and parallel execution can still make two runs differ.

The dials. The Ollama entry for qwen3.8:27b lists the parameters it applies: a temperature, a min_p, a presence penalty, a repeat penalty and a draft count of four. Every runtime exposes some version of these. Temperature and the p-cutoffs shape how the next token is picked from the prediction, the penalties discourage repetition, and the context length setting decides how much of the window the runtime actually reserves. The draft count is speculative decoding: the runtime guesses several tokens ahead and checks them in one pass. A model shipped with MTP heads, extra output layers trained to guess several tokens ahead (4.2), can do the guessing itself with no second model, and this one is. Runtimes also carry toggles for the attention kernel and for quantising the context cache; their documentation names them.

Same word, different thing: parameters. A parameter is one of the model's learned numbers, which is what 27B counts. The parameters a runtime lists are its settings: temperature, min_p, the penalties. One set is in the file and fixed; the other is on the dial and yours. A card that says "parameters" means the first, and a runtime page almost always means the second.

Two speeds. Time to first token is usually dominated by the prefill: the whole prompt read at once, compute-heavy, plus whatever queueing and overhead the server adds. Tokens per second after that is the decode: one token at a time, and often limited by how fast the parameters can be read from memory rather than by arithmetic. At low concurrency, when decode is bound by memory bandwidth, a useful ceiling is memory bandwidth divided by the bytes that must be read for each generated token; batching reuses those weight reads across more work, raising aggregate throughput rather than making one request proportionally faster. The point here is that a model can be fast at one and slow at the other, and a benchmark quoting only one is hiding the other.

One request, two speeds

time to first token
tokens per second

The two speeds move independently.

  • prefillthe whole prompt read at once, compute-heavytime to first token is usually dominated by it, plus whatever queueing and overhead the server adds
  • decodeone token at a time, often limited by how fast the parameters can be read from memory rather than by arithmetictokens per second, after the first one
A model can be fast at one and slow at the other, so a benchmark quoting only one is hiding the other.
3.3

Image generation: why does it ship with text encoders inside?

The FLUX.2 Black Forest Labs [klein] 4B card says the model fits in about 13 GB of VRAM and runs on an RTX 3090 or 4070 and above. The FLUX.2 [dev] card describes a 32-billion-parameter rectified flow transformer. Same family, two months apart, an eight-fold gap in size, and the smaller one has the number you can act on.

How it works. An image model starts from noise and removes it in steps, steered by a text description. The steering needs the text turned into numbers first, so the folder holds one or more text encoders, commonly from the CLIP or T5 families; Stable Diffusion 3.5's Stability card lists three, two CLIP encoders and a T5. These are encoders, not chat models, which is one reason prompting them differs. The image is built in a compressed space and decoded to pixels at the end by a third component, the VAE, a small model that squeezes a picture into a compact form and expands it again.

Same word, different thing: step. Here a step is one denoising pass at generation time. In a training log a step is one update of the weights. Same word, different clock.

The words. Steps are denoising passes; more is slower and, up to a point, better. Guidance (CFG on the older cards) is how hard the model is pushed towards the prompt. The sampler or scheduler is the update rule that takes each step, and the noise or timestep schedule decides which noise levels the steps visit; tools name them separately and the two are often confused. A seed makes a run reproducible under the same deterministic setup. Negative prompt is what to steer away from. Inpainting fills a masked region, outpainting extends the canvas, image-to-image starts from a picture rather than noise, and an edit model takes an instruction about an existing image; on the hub, text-to-image and image-to-image are the task tags. ControlNet and its relatives condition the image on a pose, a depth map or an edge drawing. A LoRA here is a style or a character. A turbo or distilled checkpoint is one trained to need far fewer steps (the Apex tier has a box on the other thing "distilled" means): Z-Image-Turbo's card claims eight function evaluations and sub-second latency. An upscaler is a separate model.

The architecture words on the cards are MMDiT (Stable Diffusion 3.5's card), rectified flow transformer (FLUX.2's), and MoE (HunyuanImage 3.0's card: 64 experts, 80 billion parameters, 13 billion active). You do not need any of them to choose; DiT and flow matching are the family names, and the Apex tier lists them with the rest.

The anchors. Parameters as counted in the safetensors files; the hub's diffusers library tag on every one except HunyuanImage 3.0 (transformers) and Sana (sana).

Model Parameters Licence Gated Card says
FLUX.2 [dev] 32.2B "other", flux-non-commercial-license automatic 32B rectified flow transformer; generate, edit, combine
FLUX.2 [klein] 4B 3.9B Apache 2.0 no ~13 GB VRAM; RTX 3090 / 4070 and above
Stable Diffusion 3.5 Large 8.1B "other", stabilityai-ai-community automatic MMDiT
Stable Diffusion 3.5 Medium 2.5B "other", stabilityai-ai-community automatic
SDXL base 1.0 2.6B openrail++ no 2023
Qwen-Image-2512 20.4B Apache 2.0 no improved text rendering over Qwen-Image
Z-Image-Turbo 6.2B Apache 2.0 no distilled; 8 NFEs; sub-second
HunyuanImage 3.0 83.0B "other", tencent-hunyuan-community no 64 experts, 13B active

What runs it. The diffusers library is the reference runtime and is what most cards give code for. ComfyUI is the node-based front end that the FLUX.2 [dev], Stable Diffusion 3.5, LTX-2.5 and Wan 2.2 Alibaba cards point to by name, and it is where LoRAs, ControlNets and upscalers get wired together.

Memory floor. From the cards: FLUX.2 [klein] 4B at about 13 GB. The other cards in the table do not state a figure, and the 32B and 83B models plainly need more than a consumer card without offloading. The rule is the same as for language models, parameters times bytes, plus working space that grows with resolution rather than context.

The licence trap. Look at the table. Every Stability and Black Forest Labs "dev" checkpoint is under a custom licence with automatic gating; the Apache 2.0 checkpoints are the Chinese labs' and Black Forest Labs' small ones. "Non-commercial" in a licence name means what it says, and the commercial licence is a separate purchase.

What people get wrong. Prompting an image model like a chatbot. It does not converse. Classic text-to-image prompts are descriptions; edit models are often instruction-following; either way the card's examples show the form its text encoders were trained on. Read them before writing your own.

From a sentence to pixels

your prompt
text
text encoders
turn it into numbers
denoise
in a compressed space
VAE
decodes to pixels

It starts from noise and removes it in steps, steered by the description. The image is built small and decoded at the end.

The text encoders inside

CLIPCLIPT5

The steering needs the text turned into numbers first, so the folder holds one or more text encoders, commonly from the CLIP or T5 families. Stable Diffusion 3.5's card lists these three.

Encoders, not chat models, which is one reason prompting them differs. An image model does not converse.

The two dials on the loop

  • stepsdenoising passes: more is slower and, up to a point, better
  • guidancehow hard the model is pushed towards the prompt (CFG on older cards)

A turbo or distilled checkpoint is trained to need far fewer of them: Z-Image-Turbo's card claims eight function evaluations and sub-second latency.

Classic text-to-image prompts are descriptions; edit models are often instruction-following. Either way the card's examples show the form its text encoders were trained on.
3.4

Video generation: why does the licence mention revenue?

The LTX-2.5 card offers the model at no cost for commercial use to organisations under ten million dollars in annual revenue and points larger ones to a separate agreement. Mochi 1's card gives three memory figures for one model: 60 GB on a single card, 42 GB for the highest-quality example, 22 GB with a lower-precision variant and a small drop in quality. Both facts are typical of the family.

How it works. The same denoising as image generation, with time as an extra dimension: the model generates a block of frames together so they agree with each other. The compressed space is bigger by the number of frames, which is why memory floors sit an order of magnitude above image models. The newest models generate the soundtrack in the same pass: LTX-2.5 lists text-to-audio, audio-to-video and video-to-audio among its modes.

The words. Text-to-video, image-to-video (animate a still) and video-to-video are the task tags. Reference-to-video conditions on a subject. Temporal coherence is whether the frames agree. Resolution and length limits are on the card and are the numbers that matter: Wan 2.2's TI2V-5B card says 720p at 24 frames per second, and its 14B models list 480p and 720p. Wan 2.2 also puts a mixture of experts inside a video diffusion model, which its card says was borrowed from language models.

The anchors.

Model Parameters Licence Gated Card says
LTX-2.5 22B per its repository names "other", ltx-2.x-community-license-agreement automatic audio modes; free commercial use under $10M revenue
LTX-2 18.9B "other", ltx-2-community-license-agreement no
Wan 2.2 T2V-A14B MoE, 14B active Apache 2.0 no needs a GPU with at least 80 GB VRAM as shipped
Wan 2.2 TI2V-5B 5B Apache 2.0 no runs on a 24 GB GPU such as an RTX 4090; 720p 24 fps
HunyuanVideo 1.5 not in safetensors metadata "other", tencent-hunyuan-community no minimum 14 GB with offloading
Mochi 1 preview 10.0B Apache 2.0 no 22 to 60 GB depending on precision and setup
CogVideoX-5b 5.6B "other" no 2024
Step-Video-T2V 29.3B MIT no

What runs it. The diffusers library on the LTX Lightricks, Mochi and CogVideoX cards; the makers' own code on Wan (wan2.2) and HunyuanVideo; ComfyUI on the LTX and Wan cards.

Memory floor. The cards say it themselves, and the numbers above are theirs: 14 GB minimum for HunyuanVideo 1.5 with offloading, 24 GB for the small Wan, 80 GB for the large one as shipped, 22 to 60 GB for Mochi. Offloading, moving parts of the model to system memory between steps, is what makes the low figures possible and is why the same model has three.

The licence trap. Revenue caps. The LTX-2.x community licence is the clearest case, but the pattern is a "community" licence whose terms depend on who you are. Read the card's licence section, not the badge.

What people get wrong. Assuming the successor will ship on the same terms. Between LTX-2 and LTX-2.5 the licence name changed and gating went from none to automatic; the family name stayed. Anchor on the checkpoint you have, and read the next one's card as if it were a stranger's.

3.5

Speech and audio: why is it never one model?

Whisper OpenAI large-v3 hears thirty seconds at a time; its card says so, and anything longer is cut into thirty-second windows and stitched. It transcribes and translates, lists 99 language codes on its card, and does not say who was speaking. For that you download a second model, and for deciding where speech starts and stops, a third.

How it works. A listening model usually converts audio into time-frequency or learned features and reads them the way a language model reads tokens, emitting text; Whisper's card describes its spectrogram input, a picture of which frequencies are present at each moment, and other architectures differ. A speaking model runs the other way: it predicts an acoustic or latent representation, a compressed internal form, that a decoder or vocoder turns into waveform audio. Diarisation is a separate model that clusters the audio by voice and labels the segments; it does not know what was said.

The words. ASR and STT both mean speech to text; the hub's task tag is automatic-speech-recognition. WER, word error rate, is its number. Streaming means transcribing as the audio arrives; batch means after. Real-time factor is how long it takes per second of audio. TTS is text to speech; its task tag is text-to-speech, it has no single number (listening-test scores such as MOS, intelligibility and speaker-similarity metrics are all used), and voice cloning is producing a voice from a sample (the Qwen3-TTS card says three seconds of audio is enough for its base model). Voice design is producing a voice from a description. VAD, voice activity detection, is the gatekeeper that tells the others when there is speech at all. Conversational speech models take both text and audio as input and generate speech that fits the conversation, the CSM card being the reference. Inverse text normalisation turns "two hundred" into "200" and is on the card as a feature when it is there.

Same word, different thing: batch. Three of them. In 2.2 a batch is how many requests a runtime is serving at once, which is a count. Here it is the opposite of streaming: transcribing after the audio has arrived rather than as it does, which is a timing. In 4.6 it is a file of requests sent together and returned within hours at a reduced rate, which is a price. Nothing on a page tells you which; the sentence around it does.

The anchors.

Model Job Parameters Licence Gated Card says
Whisper large-v3 ASR and translation 1.54B Apache 2.0 no >5M hours of training audio; 30-second window; 99 language codes
Whisper large-v3-turbo ASR 809M MIT no
Parakeet TDT 0.6B v3 ASR 627M CC BY 4.0 no multilingual, high throughput, punctuation and capitalisation, word timestamps
Canary 1B v2 ASR and translation 979M CC BY 4.0 no 25 European languages
Qwen3-ASR 1.7B ASR and language ID 2.35B Apache 2.0 no 30 languages and 22 Chinese dialects; vLLM-based inference toolkit
Moonshine base ASR 62M MIT no
pyannote speaker-diarization-community-1 diarisation not stated CC BY 4.0 automatic
pyannote segmentation-3.0 VAD and segmentation not stated MIT automatic
NVIDIA Sortformer 4spk v1 diarisation 124M CC BY-NC 4.0 no
NVIDIA Nemotron 3 Diarization preview diarisation not stated evaluation licence manual published 18 September 2026
Kokoro-82M TTS 82M Apache 2.0 no 82M parameters
Qwen3-TTS 1.7B Base TTS and cloning 1.93B Apache 2.0 no 10 languages; 3-second clone; VoiceDesign and CustomVoice siblings
Dia 1.6B TTS 1.61B Apache 2.0 no
CSM 1B conversational speech 1.55B Apache 2.0 automatic text and audio in; Llama backbone
PocketTTS TTS 100M per card CC BY 4.0 automatic runs on CPU
XTTS v2 TTS and cloning not stated Coqui public model licence no
F5-TTS TTS not stated CC BY-NC 4.0 no
Fish Speech 1.5 TTS not stated CC BY-NC-SA 4.0 no
Higgs Audio v2 3B TTS 5.77B in file "other" no
Stable Audio Open 1.0 text to audio 1.21B Stable Audio community automatic
Magenta RealTime 2 music not stated CC BY 4.0 no live, continuous, text-steered

What runs it. Transformers for the Whisper family and Moonshine, NVIDIA's NeMo for Parakeet, Canary and Sortformer, the pyannote-audio library for pyannote, and each TTS model's own library or code; the library field on the card says which. Qwen3-ASR ships its own vLLM-based toolkit.

Memory floor. Most listening models are small: the largest listening model above is 2.35 billion parameters and most are under one billion, so a few gigabytes of memory or a CPU is enough for one stream and the cost is in throughput, not fit. Speech generation ranges wider; the table runs from 82 million to 5.77 billion parameters, and the large end needs a GPU.

The licence trap. The Creative Commons family. CC BY 4.0 is permissive with attribution; CC BY-NC and CC BY-NC-SA are non-commercial, and three of the popular speech models above carry them. A speech pipeline is three or four models, and the deployment has to comply with every licence in it; one non-commercial component constrains the whole use case.

What people get wrong. Expecting one download to transcribe, diarise, detect speech and stream. Whisper itself does not natively provide diarisation, voice activity detection or streaming; wrappers add chunked streaming around it, and the other two are separate models. Build the pipeline, then read every licence in it.

One download

Whisper large-v31.54B · Apache 2.0transcribes and translates · 99 language codes · thirty seconds at a time

Anything longer is cut into thirty-second windows and stitched.

What it does not natively provide

  • who was speaking
  • where speech starts and stops
  • streaming; wrappers add chunked streaming around it

So the pipeline is three models

  1. 1VAD and segmentationpyannote segmentation-3.0→ where speech is at allMIT
  2. 2ASR and translationWhisper large-v3→ the wordsApache 2.0
  3. 3diarisationpyannote speaker-diarization-community-1→ who spoke, not whatCC BY 4.0

The first is the gatekeeper: it tells the others when there is speech at all.

Every licence in it counts

Of the 21 models in this section's table, 3 are non-commercial:

  • NVIDIA Sortformer 4spk v1diarisationCC BY-NC 4.0
  • F5-TTSTTSCC BY-NC 4.0
  • Fish Speech 1.5TTSCC BY-NC-SA 4.0

One of them anywhere in the pipeline constrains the whole use case.

Build the pipeline, then read every licence in it.
3.6

Vision understanding: what can a language model see?

The Qwen3-VL-32B card claims the model can operate PC and mobile interfaces, recognise elements and complete tasks, and ground objects in two and three dimensions. The SAM 3 Meta card claims the model can detect, segment and track objects in images and video from a text prompt or a point, box or mask. They are answering different questions with the same word, vision.

How it works. Understanding models come in two shapes. The older shape is task-specific: an image goes in, a label, a set of boxes, a mask or a string of text comes out, and the task tag on the hub names which (image-classification, mask-generation, zero-shot-image-classification). The newer shape is a language model with a vision encoder attached, the VLM: image and text go in, text comes out, task tag image-text-to-text, and the same model reads a document, describes a scene and answers questions about a chart.

The words. Classification names the image. Detection draws boxes. Segmentation draws masks; promptable segmentation lets you say which object. OCR reads text out of pixels; document understanding reads structure too. ViT is the vision transformer, the image-side architecture most of these share. CLIP and its successors (SigLIP Google DeepMind is one) are pairs of encoders, one for images and one for text, trained so that matching pairs land close; the CLIP card describes the contrastive training, and that pairing is what gives zero-shot classification, naming a category the model was never trained on. Grounding is pointing at where in the image something is.

The anchors.

Model Task tag Parameters Licence Gated Card says
Qwen3-VL 8B / 32B Instruct image-text-to-text 8.8B / 33.4B Apache 2.0 no GUI operation, 2D and 3D grounding
SAM 3 mask-generation 860M "other" manual promptable segmentation, text or visual prompts, images and video
SAM 2.1 hiera large mask-generation 224M Apache 2.0 no
CLIP ViT-L/14 zero-shot-image-classification 428M none stated no contrastive image and text encoders
SigLIP 2 so400m zero-shot-image-classification 1.14B Apache 2.0 no
ViT base patch16 image-classification 87M Apache 2.0 no
Florence-2 large image-text-to-text 777M MIT no
DeepSeek-OCR image-text-to-text 3.34B MIT no
PaddleOCR-VL image-text-to-text 959M Apache 2.0 no

The Gemma 4 and Qwen3.8 language models from the first section are also on this list by task tag. A VLM is its own checkpoint, with vision components inside it, and runs on the language-model runtime family; the task tag is how you spot one.

What runs it. Transformers for all of the above except PaddleOCR-VL, which lists its own library. The VLMs run on the same engines as language models, vLLM and llama.cpp included, because that is what they are.

Memory floor. The task-specific models are small, from 87 million to about a billion parameters. The VLMs are language-model sized and follow the language-model rule: the file plus the context, and images cost tokens.

The licence trap. SAM 3 is under a custom licence with manual approval while its predecessor is Apache 2.0. CLIP's OpenAI card states no licence field at all, which is not the same as permission.

What people get wrong. Using a VLM for pixel-precise work. It describes and it grounds; it does not draw a mask to the pixel. For that, the task-specific models exist, and they are a tenth the size.

One image, four readings

  • a label image-classification zero-shot-image-classification Classification names the imagezero-shot classification, naming a category the model was never trained on
  • a set of boxes Detection draws boxesGrounding is pointing at where in the image something is
  • a mask mask-generation Segmentation draws maskspromptable segmentation lets you say which object
  • a string of text OCR reads text out of pixelsdocument understanding reads structure too

an image goes in, a label, a set of boxes, a mask or a string of text comes out, and the task tag on the hub names which: two of the three tags the section lists name a label, so both sit on the first reading. The grid stands for an image and is not a picture of anything, because the section names no example and the figure will not invent one.

Understanding models come in two shapes

  • the older shapetask-specifican imageone of the four aboveThe task-specific models are small, from 87 million to about a billion parameters
  • the newer shapea language model with a vision encoder attached, the VLMimage and text go intext comes out image-text-to-textthe same model reads a document, describes a scene and answers questions about a chart

Inside both: ViT is the vision transformer, the image-side architecture most of these share. CLIP and its successors are pairs of encoders, one for images and one for text, trained so that matching pairs land close, which is what gives the zero-shot naming above.

They are answering different questions with the same word, vision. It describes and it grounds; it does not draw a mask to the pixel; for that, the task-specific models exist, and they are a tenth the size.
3.7

Embeddings and rerankers: why does the small file matter so much?

all-MiniLM-L6-v2 is 23 million parameters, Apache 2.0, and on the day this guide was checked showed a download count in the hundreds of millions, more than any language model in this guide. It generates nothing. It turns a sentence into 384 numbers.

How it works. An embedding model reads text and outputs a fixed-length list of numbers, a vector, positioned so that texts with similar meaning land near each other. Search is then geometry: embed the query, find the stored vectors closest to it. A reranker is the second stage, a model that reads the query and a candidate together and scores how well they match; it is slower and more accurate, so it runs on the shortlist the embedding search produced. This two-stage search is what the Apex tier calls retrieval, and RAG is a language model reading its results.

The words. Dimensions is the length of the vector: 384 for MiniLM, up to 1,024 for Qwen3-Embedding-0.6B, 768 for EmbeddingGemma. Matryoshka (MRL on the cards) means the vector can be truncated to a shorter one and still work; EmbeddingGemma's card lists 512, 256 and 128, and Qwen's says any length from 32 to 1,024. Cosine similarity is the usual measure of "near". Bi-encoder is the embedding shape (query and document encoded separately); cross-encoder is the reranker shape (encoded together). Semantic search is search by meaning; keyword search is by exact terms; hybrid is both, merged. Multimodal embeddings put images and text in the same space; on the hub the task tag visual-document-retrieval marks the ones built for pages and screenshots. MTEB is the benchmark: eight task types, 58 datasets and 112 languages in the paper that defined it, with a public leaderboard on the hub.

The anchors.

Model Task tag Parameters Licence Gated Card says
Qwen3-Embedding 0.6B / 4B / 8B feature-extraction 596M / 4.0B / 7.6B Apache 2.0 no dims 32 to 1,024 on the 0.6B; 100+ languages; 8B claimed top of MTEB multilingual
Qwen3-Reranker 0.6B text-ranking 596M Apache 2.0 no
Qwen3-VL-Embedding 2B sentence-similarity 2.1B Apache 2.0 no multimodal, built on Qwen3-VL
EmbeddingGemma 300M sentence-similarity 303M Gemma licence manual 768 dims, MRL to 128; on-device focus; 100+ languages in training
jina-embeddings-v5 text small feature-extraction 596M CC BY-NC 4.0 no
jina-reranker-v3 text-ranking 597M CC BY-NC 4.0 no
bge-m3 sentence-similarity not stated MIT no
nomic-embed-text-v2-moe sentence-similarity 475M Apache 2.0 no a mixture of experts in an embedding model
ColPali v1.3 visual-document-retrieval not stated MIT no built on PaliGemma
all-MiniLM-L6-v2 sentence-similarity 23M Apache 2.0 no 384 dimensions
mxbai-embed-large-v1 feature-extraction 335M Apache 2.0 no

What runs it. The sentence-transformers library is the library field on most of these cards and is the reference runtime. Several also run under the language runtimes in an embedding mode; the card's usage section says which.

Memory floor. Near zero. Everything above except the Qwen 4B and 8B fits in a couple of gigabytes, and CPU is a normal place to run them.

The licence trap. Two of the best-known families ship under CC BY-NC (Jina) and a custom gated licence (EmbeddingGemma). The embedding model is the one component you cannot swap later without re-indexing everything, so its licence is the one to settle first.

What people get wrong. Changing the embedding model between index time and query time. Every vector in the index (the stored collection of vectors, not the index file of 1.3, which lists tensors) was made by one model; a different model's query vector lands in a different space and matches nothing. Pin the model, and re-index when you change it.

Stage one · search is geometry

An embedding model turns text into a fixed-length list of numbers, positioned so that texts with similar meaning land near each other. Embed the query, then find the stored vectors closest to it.

Cosine similarity is the usual measure of "near", and a real space has far more than the two directions drawn here.

  • 384all-MiniLM-L6-v2
  • 768EmbeddingGemma
  • up to 1,024Qwen3-Embedding-0.6B

Stage two · the shortlist, read again

  • bi-encoderthe embedding shape, query and document encoded separatelyfast enough to run over everything stored
  • cross-encoderthe reranker shape, query and candidate encoded togetherslower and more accurate, so it runs only on the shortlist the first stage produced

And the way it breaks

Every vector in the index was made by one model. A different model's query vector lands in a different space and matches nothing. Pin the model, and re-index when you change it.

Which is why the embedding model is the one component you cannot swap later without re-indexing everything.

It generates nothing. That is the point of it, and why a 23-million-parameter file showed a download count, on the day the guide was checked, higher than any language model in it.
3.8

Classical machine learning: where is the AI you already have?

The scikit-learn user guide's table of contents is a map of the machine learning that ran the world before 2020 and still runs most of it: linear models, support vector machines, nearest neighbours, decision trees, ensembles of trees (gradient boosting, random forests), clustering, dimensionality reduction, outlier detection. Almost none of it is distributed as a model card on the hub, and it is less often what current AI marketing means by AI, though classifiers, recommenders and vision systems are AI by any older definition.

How it works. A classifier sorts an input into categories. A regressor predicts a number. A clustering method groups things without labels. An anomaly detector flags what does not fit. A recommender ranks items for a person. A forecaster predicts the next values of a series. Each is trained on a table of examples, produces a model that is usually kilobytes to megabytes, and runs on an ordinary processor fast enough that nobody measures it in tokens per second.

The words. Supervised means the examples had labels; unsupervised means they did not; semi-supervised is in between; the scikit-learn guide is organised by exactly that split. Features are the columns. Gradient-boosted trees are the workhorse for tabular data. Overfitting is memorising the examples rather than learning the pattern. Reinforcement learning is a way of training, not a kind of model, and the Apex tier places it.

Why it is here. Three reasons. It remains widespread in production. None of the VRAM story applies to it. And it is the anchor for a reader whose own organisation already runs it under some other name: fraud scoring, churn prediction, demand forecasting.

What people get wrong. Reaching for a language model to sort a table. If the inputs are columns and the answers are categories, the model that does it is in the scikit-learn guide, runs on a laptop, and does not hallucinate in the generative sense; it can still be confidently wrong, which is what its evaluation metrics measure.

Six shapes, and what each one does

  • A classifiersorts an input into categories
  • A regressorpredicts a number
  • A clustering methodgroups things without labels
  • An anomaly detectorflags what does not fit
  • A recommenderranks items for a person
  • A forecasterpredicts the next values of a series

Almost none of it is distributed as a model card on the hub (almost is not the same as never), and it is less often what current AI marketing means by AI.

What all six have in common

  • trained on a table of examples
  • kilobytes to megabytes
  • runs on an ordinary processor

Kilobytes to megabytes, against gigabytes for this tier's checkpoints. Fast enough that nobody measures it in tokens per second, and none of the VRAM story applies to it.

The mistake, and the test that avoids it

ifthe inputs are columnsthe answers are categories
thenthe model that does it is in the scikit-learn guide, runs on a laptop

the mistake reaching for a language model to sort a table

Such a model does not hallucinate in the generative sense, but it can still be confidently wrong, which is what its evaluation metrics measure.

Three words, and one that is not a kind of model

  • Featuresthe columns
  • Gradient-boosted treesthe workhorse for tabular data
  • Overfittingmemorising the examples rather than learning the pattern
  • Reinforcement learninga way of training, not a kind of model: the Apex tier places it

And the reason to know all this: your own organisation probably already runs it under some other name: fraud scoring, churn prediction, demand forecasting.

Columns in and categories out is not a job for a language model, and the thing that does it fits in a few megabytes.
3.9

The families, third column

The families from this tier, with what the folder holds, what opens it, the memory it needs, and the checkpoint that anchors it on the day this guide was checked.

Family Task tags on the hub Format and library What opens it Memory floor Anchor
Language text-generation, image-text-to-text safetensors (transformers); GGUF; MLX Ollama, LM Studio, llama.cpp; vLLM, SGLang for serving file plus context; 17 GB for the 27B at Q4 Qwen3.8-27B
Image text-to-image, image-to-image safetensors (diffusers) ComfyUI; Diffusers ~13 GB for a 4B (card) FLUX.2 klein 4B
Video text-to-video, image-to-video safetensors (diffusers or the maker's) ComfyUI; maker's code 14 to 80 GB (cards) Wan 2.2 TI2V-5B
Speech in automatic-speech-recognition safetensors (transformers, nemo) Transformers; NeMo a few GB or CPU Whisper large-v3
Speech out text-to-speech model's own model's own a few GB or CPU Kokoro-82M
Vision image-classification, mask-generation, image-text-to-text safetensors (transformers) Transformers; the language runtimes for VLMs 87M-param models on CPU; VLMs as language models SAM 3; Qwen3-VL
Embeddings sentence-similarity, feature-extraction, text-ranking safetensors (sentence-transformers) sentence-transformers; vLLM; Ollama near zero all-MiniLM-L6-v2; Qwen3-Embedding
Classical none; not on the hub pickle, ONNX, the library's own scikit-learn and relatives CPU the scikit-learn user guide

3.10The question Nexus leaves you with

You can choose a model for a job. What you cannot yet read is the release post that announced it: the training words, the architecture words, the words on the invoice, and the things bolted on around it. That is the Apex tier.