The Nexus
For the person about to download one, who wants to know which, and what it will need.
Every family in this tier gets the same treatment: what it does, how it works in a few sentences, the words on its cards, the models that anchor it on the day this guide was checked, what runs it, the memory it needs, the licence trap, and the one thing people get wrong. The anchors are a snapshot, and the date matters more here than anywhere else in the guide.
Language models: which variant are you looking at?
Gemma 4 ships in five sizes, and each size ships as a base model and an instruction-tuned one; Qwen3.8-27B ships with thinking on by default and a switch to turn it off; gpt-oss OpenAI ships with a reasoning-effort dial marked low, medium and high. Before you have chosen a family you are already choosing among variants.
You have met this: the picker, again. A chat product's picker is this choice with the filenames taken off: often a fast model and a thinking one, sometimes a small one and a large one. The card's words are the picker's words, and the rest of this section is what they mean. A coding agent may expose the same sort of choice; a coder variant, where a family ships one, is a separate checkpoint and not a setting.
How it works. The loop the Origin tier gave you: predict the next token from all the tokens so far, pick one, append it, ask again, until the model emits its stop token. A chat is that loop, wrapped in the template from the Vector tier.
The words. Two paragraphs here, the variants and the context; the next section takes what the model emits, the dials and the two speeds.
The variants. A base model is the raw result of pre-training; it continues text and is not built to answer. An instruct or chat model has been post-trained to answer; the -it on Gemma's Google DeepMind names and -Instruct on Llama's Meta mean this, and Qwen's plain Qwen3.8-27B is already post-trained according to its card. A thinking or reasoning model spends tokens on working before its answer, which may be shown, hidden or summarised depending on the system; Qwen's Alibaba card calls it thinking mode and makes it default, gpt-oss calls it reasoning effort, and Gemma 4 offers a token budget for it in five fixed steps from 70 to 1,120 tokens. That working is output tokens, as the Vector tier said, and gpt-oss's card is explicit that it is not meant to be shown to end users. Vision or VL variants take images; on the hub the task tag image-text-to-text marks them, and Qwen3.8-27B, Gemma 4 and both Llama 4 models carry it. Coder variants exist as separate checkpoints in most families.
Context. The window is a token count from the card: 262,144 for Qwen3.8-27B and the larger Gemma 4 models, 131,072 for the Gemma 4 edge models and gpt-oss, 1,048,576 for DeepSeek V4 and Kimi K3 Moonshot; GLM-5.3's config exposes the same figure and its card evaluates at 1M without stating a headline context length. It is a limit on what the model can see, not on what it knows, and a bigger window costs memory for the cache that holds it. That cache, the KV cache, has its own guide; here it is enough to know it exists and grows with context. It is also request state: each active request needs cache space of its own, apart from any prefix the runtime can share, so the weights stay fixed while the cache grows with every request being served at once, and a server can become limited by cache space even though the weights are shared. vLLM prints both when it starts, the cache's capacity in tokens and an estimate of how many requests of a given length it can serve at once.
The anchors. Licence and gating as read from the hub on the day this guide was checked. Total parameters as counted in the safetensors files (the file), with the card's own total where it differs; active parameters and routing from the card; context as the maker advertises it on the card, which is the supported figure, since a config can expose a larger number than the maker will stand behind.
| Model | Total (file) | Active (card) | Experts and routing | Context (card) | Licence | Gated |
|---|---|---|---|---|---|---|
| Qwen3.8-27B | 27.8B | dense | none | 262,144 native, 1M extended | Apache 2.0 | no |
| Qwen3.8-Flash-Next | 180B (125B model, 51B n-gram embedding, 4B MTP) | 6B | 512 experts; 10 routed + 1 shared | 262,144 | "other", qwen-community-1.0 | no |
| DeepSeek-V4.1-Flash | 763B (552B backbone, 196B Engram memory) | 8B / 16B | 384 experts; 6 per token | 1,048,576 | MIT | no |
| DeepSeek-V4-Pro-0813 | 1,651B (1.6T backbone) | 49B per the V4.1 card's table | 384 experts; 6 per token | 1,048,576 | MIT | no |
| Kimi-K3 | 2,780B (2.8T) | 104B | 896 experts; 16 selected + 2 shared | 1,048,576 | "other", kimi-k3 | no |
| GLM-5.3 | 753B | not stated on card | 256 experts; 8 per token | 1M used in card evaluations; 1,048,576 in config | "other", glm-5.3 | no |
| Llama 4 Scout | 109B | 17B | 16 experts | 10M | "other", llama4 | manual |
| Llama 4 Maverick | 402B (400B) | 17B | 128 experts | 1M | "other", llama4 | manual |
| Gemma 4 26B-A4B | 25.8B (25.2B) | 3.8B | 128 experts; 8 active + 1 shared | 256K | Apache 2.0 | no |
| Gemma 4 12B | 12.0B | dense | none | 256K | Apache 2.0 | no |
| Gemma 4 E4B / E2B | 8.0B / 5.1B | 4.5B / 2.3B "effective" | none | 128K | Apache 2.0 | no |
| gpt-oss-120b / 20b | 117B / 21B | 5.1B / 3.6B | MoE, per card | 131,072 | Apache 2.0 | no |
| Mistral Small 4 | 119B | 6.5B (branded A6B) | 128 experts; 4 active | 256K | Apache 2.0 | no |
| Phi-4 | 14.7B | dense | none | 16K | MIT | no |
| Nemotron 3 Nano 30B-A3B | 31.6B | 3B per its name | 128 experts; 6 per token | 262,144 | "other", nvidia-nemotron-open-model-license | no |
Three things in that table to notice. Almost every large model in this table is a mixture of experts, and the active counts are small next to the totals: Kimi's 2.8 trillion runs 104 billion per token, 16 of 896 experts plus 2 shared. The million-token context sits at the top of the table and Llama 4 Scout claims ten million. And the Gemma edge models carry two parameter counts, because the "E" means effective: the card says the E2B file holds 5.1 billion parameters of which 2.3 billion are counted as effective, the rest being per-layer embeddings kept out of the count. The trade's word for the small end of the table is small language model, or SLM, and it is a size, not a kind.
What runs it. The Qwen card names Transformers, vLLM, SGLang and TokenSpeed, and says the last three are for production and high throughput. The GGUF world, llama.cpp and the tools built on it (Ollama and LM Studio among them), is optimised first for local inference on one machine, though it can serve an API to several users; vLLM and SGLang are designed around high-throughput serving and batching. Both worlds advertise an OpenAI-compatible endpoint, an address that accepts requests in the shape one provider's API made common, which means a basic chat call can move between them, or to a provider's API, with little change; tool calling, reasoning controls, image input and structured-output extensions (4.6) differ between providers, and those parts of your code will not move as cleanly.
Memory floor. The weights, plus the KV cache for the context you will use, plus the runtime's own working space; a 15.5 GB file does not comfortably fit a 16 GB card. For Qwen3.8-27B: 16.46 GB for the Q4_K_M plus 0.93 GB for the vision projector, before any context. gpt-oss's own card gives its floors directly: the 120b on a single 80 GB card, the 20b within 16 GB, both because the MoE weights were post-trained at MXFP4. MXFP4 is the four-bit floating-point member of the Microscaling (MX) formats, which pair a narrow number type with a scale shared across a block of values; gpt-oss ships its MoE weights in it from the maker, rather than as a later third-party quantisation. For everything else, read the file sizes in the quantiser's folder and add room for the window you intend to use. Where the quantiser publishes quality measurements, look for perplexity or KL divergence against the BF16 or base model; that comparison is the measured loss the Vector tier meant. When the weights do not fit on one accelerator, runtimes can split the model across several; tensor parallelism, named at 2.12, is one common way to do it.
The licence trap. The same family carries permissive and restricted checkpoints side by side. On the day, Qwen3.8-27B was Apache 2.0 and Qwen3.8-Flash-Next, released two weeks later by the same lab, was under a Qwen community licence. "Qwen is Apache" is not a sentence; "this checkpoint is Apache" is.
What people get wrong. Choosing by leaderboard rank. The rank measures the vendor's benchmarks on the vendor's settings, and a smaller model that fits your card with room for context beats a larger one that spills. Choose by fit, then by the languages you need (Gemma 4's card claims 35+ supported and 140+ in pre-training; test yours), then by the task, and only then by the table.
New here: attention, the KV cache and the Transformer. How the next-token prediction is computed, why context costs memory, and what "Gated DeltaNet" on the Qwen card means belong to neural networks, and to attention and the KV cache. Nothing here needs them.
Language models: the dials and the two speeds
What it emits. Before the dials make sense, one paragraph on what the model produces. For each position it emits a score, a logit, for every token in its vocabulary. Those scores are turned into a probability distribution. Then a token is chosen from it: greedy decoding takes the most probable one every time and is deterministic; a sampler draws from the distribution, and temperature, top-p, min_p and the penalties reshape the distribution before the draw. That draw is why the same prompt gives a different answer twice; a fixed seed improves reproducibility on the same runtime and hardware, but kernels, batching and parallel execution can still make two runs differ.
The dials. The Ollama entry for qwen3.8:27b lists the parameters it applies: a temperature, a min_p, a presence penalty, a repeat penalty and a draft count of four. Every runtime exposes some version of these. Temperature and the p-cutoffs shape how the next token is picked from the prediction, the penalties discourage repetition, and the context length setting decides how much of the window the runtime actually reserves. The draft count is speculative decoding: the runtime guesses several tokens ahead and checks them in one pass. A model shipped with MTP heads, extra output layers trained to guess several tokens ahead (4.2), can do the guessing itself with no second model, and this one is. Runtimes also carry toggles for the attention kernel and for quantising the context cache; their documentation names them.
Same word, different thing: parameters. A parameter is one of the model's learned numbers, which is what 27B counts. The parameters a runtime lists are its settings: temperature, min_p, the penalties. One set is in the file and fixed; the other is on the dial and yours. A card that says "parameters" means the first, and a runtime page almost always means the second.
Two speeds. Time to first token is usually dominated by the prefill: the whole prompt read at once, compute-heavy, plus whatever queueing and overhead the server adds. Tokens per second after that is the decode: one token at a time, and often limited by how fast the parameters can be read from memory rather than by arithmetic. At low concurrency, when decode is bound by memory bandwidth, a useful ceiling is memory bandwidth divided by the bytes that must be read for each generated token; batching reuses those weight reads across more work, raising aggregate throughput rather than making one request proportionally faster. The point here is that a model can be fast at one and slow at the other, and a benchmark quoting only one is hiding the other.
One request, two speeds
- prefillthe whole prompt read at once, compute-heavytime to first token is usually dominated by it, plus whatever queueing and overhead the server adds
- decodeone token at a time, often limited by how fast the parameters can be read from memory rather than by arithmetictokens per second, after the first one
Image generation: why does it ship with text encoders inside?
The FLUX.2 Black Forest Labs [klein] 4B card says the model fits in about 13 GB of VRAM and runs on an RTX 3090 or 4070 and above. The FLUX.2 [dev] card describes a 32-billion-parameter rectified flow transformer. Same family, two months apart, an eight-fold gap in size, and the smaller one has the number you can act on.
How it works. An image model starts from noise and removes it in steps, steered by a text description. The steering needs the text turned into numbers first, so the folder holds one or more text encoders, commonly from the CLIP or T5 families; Stable Diffusion 3.5's Stability card lists three, two CLIP encoders and a T5. These are encoders, not chat models, which is one reason prompting them differs. The image is built in a compressed space and decoded to pixels at the end by a third component, the VAE, a small model that squeezes a picture into a compact form and expands it again.
Same word, different thing: step. Here a step is one denoising pass at generation time. In a training log a step is one update of the weights. Same word, different clock.
The words. Steps are denoising passes; more is slower and, up to a point, better. Guidance (CFG on the older cards) is how hard the model is pushed towards the prompt. The sampler or scheduler is the update rule that takes each step, and the noise or timestep schedule decides which noise levels the steps visit; tools name them separately and the two are often confused. A seed makes a run reproducible under the same deterministic setup. Negative prompt is what to steer away from. Inpainting fills a masked region, outpainting extends the canvas, image-to-image starts from a picture rather than noise, and an edit model takes an instruction about an existing image; on the hub, text-to-image and image-to-image are the task tags. ControlNet and its relatives condition the image on a pose, a depth map or an edge drawing. A LoRA here is a style or a character. A turbo or distilled checkpoint is one trained to need far fewer steps (the Apex tier has a box on the other thing "distilled" means): Z-Image-Turbo's card claims eight function evaluations and sub-second latency. An upscaler is a separate model.
The architecture words on the cards are MMDiT (Stable Diffusion 3.5's card), rectified flow transformer (FLUX.2's), and MoE (HunyuanImage 3.0's card: 64 experts, 80 billion parameters, 13 billion active). You do not need any of them to choose; DiT and flow matching are the family names, and the Apex tier lists them with the rest.
The anchors. Parameters as counted in the safetensors files; the hub's diffusers library tag on every one except HunyuanImage 3.0 (transformers) and Sana (sana).
| Model | Parameters | Licence | Gated | Card says |
|---|---|---|---|---|
| FLUX.2 [dev] | 32.2B | "other", flux-non-commercial-license | automatic | 32B rectified flow transformer; generate, edit, combine |
| FLUX.2 [klein] 4B | 3.9B | Apache 2.0 | no | ~13 GB VRAM; RTX 3090 / 4070 and above |
| Stable Diffusion 3.5 Large | 8.1B | "other", stabilityai-ai-community | automatic | MMDiT |
| Stable Diffusion 3.5 Medium | 2.5B | "other", stabilityai-ai-community | automatic | |
| SDXL base 1.0 | 2.6B | openrail++ | no | 2023 |
| Qwen-Image-2512 | 20.4B | Apache 2.0 | no | improved text rendering over Qwen-Image |
| Z-Image-Turbo | 6.2B | Apache 2.0 | no | distilled; 8 NFEs; sub-second |
| HunyuanImage 3.0 | 83.0B | "other", tencent-hunyuan-community | no | 64 experts, 13B active |
What runs it. The diffusers library is the reference runtime and is what most cards give code for. ComfyUI is the node-based front end that the FLUX.2 [dev], Stable Diffusion 3.5, LTX-2.5 and Wan 2.2 Alibaba cards point to by name, and it is where LoRAs, ControlNets and upscalers get wired together.
Memory floor. From the cards: FLUX.2 [klein] 4B at about 13 GB. The other cards in the table do not state a figure, and the 32B and 83B models plainly need more than a consumer card without offloading. The rule is the same as for language models, parameters times bytes, plus working space that grows with resolution rather than context.
The licence trap. Look at the table. Every Stability and Black Forest Labs "dev" checkpoint is under a custom licence with automatic gating; the Apache 2.0 checkpoints are the Chinese labs' and Black Forest Labs' small ones. "Non-commercial" in a licence name means what it says, and the commercial licence is a separate purchase.
What people get wrong. Prompting an image model like a chatbot. It does not converse. Classic text-to-image prompts are descriptions; edit models are often instruction-following; either way the card's examples show the form its text encoders were trained on. Read them before writing your own.
From a sentence to pixels
The text encoders inside
The steering needs the text turned into numbers first, so the folder holds one or more text encoders, commonly from the CLIP or T5 families. Stable Diffusion 3.5's card lists these three.
Encoders, not chat models, which is one reason prompting them differs. An image model does not converse.
The two dials on the loop
- stepsdenoising passes: more is slower and, up to a point, better
- guidancehow hard the model is pushed towards the prompt (CFG on older cards)
Video generation: why does the licence mention revenue?
The LTX-2.5 card offers the model at no cost for commercial use to organisations under ten million dollars in annual revenue and points larger ones to a separate agreement. Mochi 1's card gives three memory figures for one model: 60 GB on a single card, 42 GB for the highest-quality example, 22 GB with a lower-precision variant and a small drop in quality. Both facts are typical of the family.
How it works. The same denoising as image generation, with time as an extra dimension: the model generates a block of frames together so they agree with each other. The compressed space is bigger by the number of frames, which is why memory floors sit an order of magnitude above image models. The newest models generate the soundtrack in the same pass: LTX-2.5 lists text-to-audio, audio-to-video and video-to-audio among its modes.
The words. Text-to-video, image-to-video (animate a still) and video-to-video are the task tags. Reference-to-video conditions on a subject. Temporal coherence is whether the frames agree. Resolution and length limits are on the card and are the numbers that matter: Wan 2.2's TI2V-5B card says 720p at 24 frames per second, and its 14B models list 480p and 720p. Wan 2.2 also puts a mixture of experts inside a video diffusion model, which its card says was borrowed from language models.
The anchors.
| Model | Parameters | Licence | Gated | Card says |
|---|---|---|---|---|
| LTX-2.5 | 22B per its repository names | "other", ltx-2.x-community-license-agreement | automatic | audio modes; free commercial use under $10M revenue |
| LTX-2 | 18.9B | "other", ltx-2-community-license-agreement | no | |
| Wan 2.2 T2V-A14B | MoE, 14B active | Apache 2.0 | no | needs a GPU with at least 80 GB VRAM as shipped |
| Wan 2.2 TI2V-5B | 5B | Apache 2.0 | no | runs on a 24 GB GPU such as an RTX 4090; 720p 24 fps |
| HunyuanVideo 1.5 | not in safetensors metadata | "other", tencent-hunyuan-community | no | minimum 14 GB with offloading |
| Mochi 1 preview | 10.0B | Apache 2.0 | no | 22 to 60 GB depending on precision and setup |
| CogVideoX-5b | 5.6B | "other" | no | 2024 |
| Step-Video-T2V | 29.3B | MIT | no |
What runs it. The diffusers library on the LTX Lightricks, Mochi and CogVideoX cards; the makers' own code on Wan (wan2.2) and HunyuanVideo; ComfyUI on the LTX and Wan cards.
Memory floor. The cards say it themselves, and the numbers above are theirs: 14 GB minimum for HunyuanVideo 1.5 with offloading, 24 GB for the small Wan, 80 GB for the large one as shipped, 22 to 60 GB for Mochi. Offloading, moving parts of the model to system memory between steps, is what makes the low figures possible and is why the same model has three.
The licence trap. Revenue caps. The LTX-2.x community licence is the clearest case, but the pattern is a "community" licence whose terms depend on who you are. Read the card's licence section, not the badge.
What people get wrong. Assuming the successor will ship on the same terms. Between LTX-2 and LTX-2.5 the licence name changed and gating went from none to automatic; the family name stayed. Anchor on the checkpoint you have, and read the next one's card as if it were a stranger's.
Speech and audio: why is it never one model?
Whisper OpenAI large-v3 hears thirty seconds at a time; its card says so, and anything longer is cut into thirty-second windows and stitched. It transcribes and translates, lists 99 language codes on its card, and does not say who was speaking. For that you download a second model, and for deciding where speech starts and stops, a third.
How it works. A listening model usually converts audio into time-frequency or learned features and reads them the way a language model reads tokens, emitting text; Whisper's card describes its spectrogram input, a picture of which frequencies are present at each moment, and other architectures differ. A speaking model runs the other way: it predicts an acoustic or latent representation, a compressed internal form, that a decoder or vocoder turns into waveform audio. Diarisation is a separate model that clusters the audio by voice and labels the segments; it does not know what was said.
The words. ASR and STT both mean speech to text; the hub's task tag is automatic-speech-recognition. WER, word error rate, is its number. Streaming means transcribing as the audio arrives; batch means after. Real-time factor is how long it takes per second of audio. TTS is text to speech; its task tag is text-to-speech, it has no single number (listening-test scores such as MOS, intelligibility and speaker-similarity metrics are all used), and voice cloning is producing a voice from a sample (the Qwen3-TTS card says three seconds of audio is enough for its base model). Voice design is producing a voice from a description. VAD, voice activity detection, is the gatekeeper that tells the others when there is speech at all. Conversational speech models take both text and audio as input and generate speech that fits the conversation, the CSM card being the reference. Inverse text normalisation turns "two hundred" into "200" and is on the card as a feature when it is there.
Same word, different thing: batch. Three of them. In 2.2 a batch is how many requests a runtime is serving at once, which is a count. Here it is the opposite of streaming: transcribing after the audio has arrived rather than as it does, which is a timing. In 4.6 it is a file of requests sent together and returned within hours at a reduced rate, which is a price. Nothing on a page tells you which; the sentence around it does.
The anchors.
| Model | Job | Parameters | Licence | Gated | Card says |
|---|---|---|---|---|---|
| Whisper large-v3 | ASR and translation | 1.54B | Apache 2.0 | no | >5M hours of training audio; 30-second window; 99 language codes |
| Whisper large-v3-turbo | ASR | 809M | MIT | no | |
| Parakeet TDT 0.6B v3 | ASR | 627M | CC BY 4.0 | no | multilingual, high throughput, punctuation and capitalisation, word timestamps |
| Canary 1B v2 | ASR and translation | 979M | CC BY 4.0 | no | 25 European languages |
| Qwen3-ASR 1.7B | ASR and language ID | 2.35B | Apache 2.0 | no | 30 languages and 22 Chinese dialects; vLLM-based inference toolkit |
| Moonshine base | ASR | 62M | MIT | no | |
| pyannote speaker-diarization-community-1 | diarisation | not stated | CC BY 4.0 | automatic | |
| pyannote segmentation-3.0 | VAD and segmentation | not stated | MIT | automatic | |
| NVIDIA Sortformer 4spk v1 | diarisation | 124M | CC BY-NC 4.0 | no | |
| NVIDIA Nemotron 3 Diarization preview | diarisation | not stated | evaluation licence | manual | published 18 September 2026 |
| Kokoro-82M | TTS | 82M | Apache 2.0 | no | 82M parameters |
| Qwen3-TTS 1.7B Base | TTS and cloning | 1.93B | Apache 2.0 | no | 10 languages; 3-second clone; VoiceDesign and CustomVoice siblings |
| Dia 1.6B | TTS | 1.61B | Apache 2.0 | no | |
| CSM 1B | conversational speech | 1.55B | Apache 2.0 | automatic | text and audio in; Llama backbone |
| PocketTTS | TTS | 100M per card | CC BY 4.0 | automatic | runs on CPU |
| XTTS v2 | TTS and cloning | not stated | Coqui public model licence | no | |
| F5-TTS | TTS | not stated | CC BY-NC 4.0 | no | |
| Fish Speech 1.5 | TTS | not stated | CC BY-NC-SA 4.0 | no | |
| Higgs Audio v2 3B | TTS | 5.77B in file | "other" | no | |
| Stable Audio Open 1.0 | text to audio | 1.21B | Stable Audio community | automatic | |
| Magenta RealTime 2 | music | not stated | CC BY 4.0 | no | live, continuous, text-steered |
What runs it. Transformers for the Whisper family and Moonshine, NVIDIA's NeMo for Parakeet, Canary and Sortformer, the pyannote-audio library for pyannote, and each TTS model's own library or code; the library field on the card says which. Qwen3-ASR ships its own vLLM-based toolkit.
Memory floor. Most listening models are small: the largest listening model above is 2.35 billion parameters and most are under one billion, so a few gigabytes of memory or a CPU is enough for one stream and the cost is in throughput, not fit. Speech generation ranges wider; the table runs from 82 million to 5.77 billion parameters, and the large end needs a GPU.
The licence trap. The Creative Commons family. CC BY 4.0 is permissive with attribution; CC BY-NC and CC BY-NC-SA are non-commercial, and three of the popular speech models above carry them. A speech pipeline is three or four models, and the deployment has to comply with every licence in it; one non-commercial component constrains the whole use case.
What people get wrong. Expecting one download to transcribe, diarise, detect speech and stream. Whisper itself does not natively provide diarisation, voice activity detection or streaming; wrappers add chunked streaming around it, and the other two are separate models. Build the pipeline, then read every licence in it.
One download
What it does not natively provide
- who was speaking
- where speech starts and stops
- streaming; wrappers add chunked streaming around it
So the pipeline is three models
- 1VAD and segmentationpyannote segmentation-3.0→ where speech is at allMIT
- 2ASR and translationWhisper large-v3→ the wordsApache 2.0
- 3diarisationpyannote speaker-diarization-community-1→ who spoke, not whatCC BY 4.0
Every licence in it counts
Of the 21 models in this section's table, 3 are non-commercial:
- NVIDIA Sortformer 4spk v1diarisationCC BY-NC 4.0
- F5-TTSTTSCC BY-NC 4.0
- Fish Speech 1.5TTSCC BY-NC-SA 4.0
One of them anywhere in the pipeline constrains the whole use case.
Vision understanding: what can a language model see?
The Qwen3-VL-32B card claims the model can operate PC and mobile interfaces, recognise elements and complete tasks, and ground objects in two and three dimensions. The SAM 3 Meta card claims the model can detect, segment and track objects in images and video from a text prompt or a point, box or mask. They are answering different questions with the same word, vision.
How it works. Understanding models come in two shapes. The older shape is task-specific: an image goes in, a label, a set of boxes, a mask or a string of text comes out, and the task tag on the hub names which (image-classification, mask-generation, zero-shot-image-classification). The newer shape is a language model with a vision encoder attached, the VLM: image and text go in, text comes out, task tag image-text-to-text, and the same model reads a document, describes a scene and answers questions about a chart.
The words. Classification names the image. Detection draws boxes. Segmentation draws masks; promptable segmentation lets you say which object. OCR reads text out of pixels; document understanding reads structure too. ViT is the vision transformer, the image-side architecture most of these share. CLIP and its successors (SigLIP Google DeepMind is one) are pairs of encoders, one for images and one for text, trained so that matching pairs land close; the CLIP card describes the contrastive training, and that pairing is what gives zero-shot classification, naming a category the model was never trained on. Grounding is pointing at where in the image something is.
The anchors.
| Model | Task tag | Parameters | Licence | Gated | Card says |
|---|---|---|---|---|---|
| Qwen3-VL 8B / 32B Instruct | image-text-to-text | 8.8B / 33.4B | Apache 2.0 | no | GUI operation, 2D and 3D grounding |
| SAM 3 | mask-generation | 860M | "other" | manual | promptable segmentation, text or visual prompts, images and video |
| SAM 2.1 hiera large | mask-generation | 224M | Apache 2.0 | no | |
| CLIP ViT-L/14 | zero-shot-image-classification | 428M | none stated | no | contrastive image and text encoders |
| SigLIP 2 so400m | zero-shot-image-classification | 1.14B | Apache 2.0 | no | |
| ViT base patch16 | image-classification | 87M | Apache 2.0 | no | |
| Florence-2 large | image-text-to-text | 777M | MIT | no | |
| DeepSeek-OCR | image-text-to-text | 3.34B | MIT | no | |
| PaddleOCR-VL | image-text-to-text | 959M | Apache 2.0 | no |
The Gemma 4 and Qwen3.8 language models from the first section are also on this list by task tag. A VLM is its own checkpoint, with vision components inside it, and runs on the language-model runtime family; the task tag is how you spot one.
What runs it. Transformers for all of the above except PaddleOCR-VL, which lists its own library. The VLMs run on the same engines as language models, vLLM and llama.cpp included, because that is what they are.
Memory floor. The task-specific models are small, from 87 million to about a billion parameters. The VLMs are language-model sized and follow the language-model rule: the file plus the context, and images cost tokens.
The licence trap. SAM 3 is under a custom licence with manual approval while its predecessor is Apache 2.0. CLIP's OpenAI card states no licence field at all, which is not the same as permission.
What people get wrong. Using a VLM for pixel-precise work. It describes and it grounds; it does not draw a mask to the pixel. For that, the task-specific models exist, and they are a tenth the size.
One image, four readings
- a label
image-classificationzero-shot-image-classificationClassification names the imagezero-shot classification, naming a category the model was never trained on - a set of boxes Detection draws boxesGrounding is pointing at where in the image something is
- a mask
mask-generationSegmentation draws maskspromptable segmentation lets you say which object - a string of text OCR reads text out of pixelsdocument understanding reads structure too
Understanding models come in two shapes
- the older shapetask-specifican imageone of the four aboveThe task-specific models are small, from 87 million to about a billion parameters
- the newer shapea language model with a vision encoder attached, the VLMimage and text go intext comes out
image-text-to-textthe same model reads a document, describes a scene and answers questions about a chart
Embeddings and rerankers: why does the small file matter so much?
all-MiniLM-L6-v2 is 23 million parameters, Apache 2.0, and on the day this guide was checked showed a download count in the hundreds of millions, more than any language model in this guide. It generates nothing. It turns a sentence into 384 numbers.
How it works. An embedding model reads text and outputs a fixed-length list of numbers, a vector, positioned so that texts with similar meaning land near each other. Search is then geometry: embed the query, find the stored vectors closest to it. A reranker is the second stage, a model that reads the query and a candidate together and scores how well they match; it is slower and more accurate, so it runs on the shortlist the embedding search produced. This two-stage search is what the Apex tier calls retrieval, and RAG is a language model reading its results.
The words. Dimensions is the length of the vector: 384 for MiniLM, up to 1,024 for Qwen3-Embedding-0.6B, 768 for EmbeddingGemma. Matryoshka (MRL on the cards) means the vector can be truncated to a shorter one and still work; EmbeddingGemma's card lists 512, 256 and 128, and Qwen's says any length from 32 to 1,024. Cosine similarity is the usual measure of "near". Bi-encoder is the embedding shape (query and document encoded separately); cross-encoder is the reranker shape (encoded together). Semantic search is search by meaning; keyword search is by exact terms; hybrid is both, merged. Multimodal embeddings put images and text in the same space; on the hub the task tag visual-document-retrieval marks the ones built for pages and screenshots. MTEB is the benchmark: eight task types, 58 datasets and 112 languages in the paper that defined it, with a public leaderboard on the hub.
The anchors.
| Model | Task tag | Parameters | Licence | Gated | Card says |
|---|---|---|---|---|---|
| Qwen3-Embedding 0.6B / 4B / 8B | feature-extraction | 596M / 4.0B / 7.6B | Apache 2.0 | no | dims 32 to 1,024 on the 0.6B; 100+ languages; 8B claimed top of MTEB multilingual |
| Qwen3-Reranker 0.6B | text-ranking | 596M | Apache 2.0 | no | |
| Qwen3-VL-Embedding 2B | sentence-similarity | 2.1B | Apache 2.0 | no | multimodal, built on Qwen3-VL |
| EmbeddingGemma 300M | sentence-similarity | 303M | Gemma licence | manual | 768 dims, MRL to 128; on-device focus; 100+ languages in training |
| jina-embeddings-v5 text small | feature-extraction | 596M | CC BY-NC 4.0 | no | |
| jina-reranker-v3 | text-ranking | 597M | CC BY-NC 4.0 | no | |
| bge-m3 | sentence-similarity | not stated | MIT | no | |
| nomic-embed-text-v2-moe | sentence-similarity | 475M | Apache 2.0 | no | a mixture of experts in an embedding model |
| ColPali v1.3 | visual-document-retrieval | not stated | MIT | no | built on PaliGemma |
| all-MiniLM-L6-v2 | sentence-similarity | 23M | Apache 2.0 | no | 384 dimensions |
| mxbai-embed-large-v1 | feature-extraction | 335M | Apache 2.0 | no |
What runs it. The sentence-transformers library is the library field on most of these cards and is the reference runtime. Several also run under the language runtimes in an embedding mode; the card's usage section says which.
Memory floor. Near zero. Everything above except the Qwen 4B and 8B fits in a couple of gigabytes, and CPU is a normal place to run them.
The licence trap. Two of the best-known families ship under CC BY-NC (Jina) and a custom gated licence (EmbeddingGemma). The embedding model is the one component you cannot swap later without re-indexing everything, so its licence is the one to settle first.
What people get wrong. Changing the embedding model between index time and query time. Every vector in the index (the stored collection of vectors, not the index file of 1.3, which lists tensors) was made by one model; a different model's query vector lands in a different space and matches nothing. Pin the model, and re-index when you change it.
Stage one · search is geometry
An embedding model turns text into a fixed-length list of numbers, positioned so that texts with similar meaning land near each other. Embed the query, then find the stored vectors closest to it.
- 384all-MiniLM-L6-v2
- 768EmbeddingGemma
- up to 1,024Qwen3-Embedding-0.6B
Stage two · the shortlist, read again
- bi-encoderthe embedding shape, query and document encoded separatelyfast enough to run over everything stored
- cross-encoderthe reranker shape, query and candidate encoded togetherslower and more accurate, so it runs only on the shortlist the first stage produced
And the way it breaks
Every vector in the index was made by one model. A different model's query vector lands in a different space and matches nothing. Pin the model, and re-index when you change it.
Classical machine learning: where is the AI you already have?
The scikit-learn user guide's table of contents is a map of the machine learning that ran the world before 2020 and still runs most of it: linear models, support vector machines, nearest neighbours, decision trees, ensembles of trees (gradient boosting, random forests), clustering, dimensionality reduction, outlier detection. Almost none of it is distributed as a model card on the hub, and it is less often what current AI marketing means by AI, though classifiers, recommenders and vision systems are AI by any older definition.
How it works. A classifier sorts an input into categories. A regressor predicts a number. A clustering method groups things without labels. An anomaly detector flags what does not fit. A recommender ranks items for a person. A forecaster predicts the next values of a series. Each is trained on a table of examples, produces a model that is usually kilobytes to megabytes, and runs on an ordinary processor fast enough that nobody measures it in tokens per second.
The words. Supervised means the examples had labels; unsupervised means they did not; semi-supervised is in between; the scikit-learn guide is organised by exactly that split. Features are the columns. Gradient-boosted trees are the workhorse for tabular data. Overfitting is memorising the examples rather than learning the pattern. Reinforcement learning is a way of training, not a kind of model, and the Apex tier places it.
Why it is here. Three reasons. It remains widespread in production. None of the VRAM story applies to it. And it is the anchor for a reader whose own organisation already runs it under some other name: fraud scoring, churn prediction, demand forecasting.
What people get wrong. Reaching for a language model to sort a table. If the inputs are columns and the answers are categories, the model that does it is in the scikit-learn guide, runs on a laptop, and does not hallucinate in the generative sense; it can still be confidently wrong, which is what its evaluation metrics measure.
Six shapes, and what each one does
- A classifiersorts an input into categories
- A regressorpredicts a number
- A clustering methodgroups things without labels
- An anomaly detectorflags what does not fit
- A recommenderranks items for a person
- A forecasterpredicts the next values of a series
What all six have in common
The mistake, and the test that avoids it
the mistake reaching for a language model to sort a table
Three words, and one that is not a kind of model
- Featuresthe columns
- Gradient-boosted treesthe workhorse for tabular data
- Overfittingmemorising the examples rather than learning the pattern
- Reinforcement learninga way of training, not a kind of model: the Apex tier places it
The families, third column
The families from this tier, with what the folder holds, what opens it, the memory it needs, and the checkpoint that anchors it on the day this guide was checked.
| Family | Task tags on the hub | Format and library | What opens it | Memory floor | Anchor |
|---|---|---|---|---|---|
| Language | text-generation, image-text-to-text | safetensors (transformers); GGUF; MLX |
Ollama, LM Studio, llama.cpp; vLLM, SGLang for serving | file plus context; 17 GB for the 27B at Q4 | Qwen3.8-27B |
| Image | text-to-image, image-to-image | safetensors (diffusers) |
ComfyUI; Diffusers | ~13 GB for a 4B (card) | FLUX.2 klein 4B |
| Video | text-to-video, image-to-video | safetensors (diffusers or the maker's) |
ComfyUI; maker's code | 14 to 80 GB (cards) | Wan 2.2 TI2V-5B |
| Speech in | automatic-speech-recognition | safetensors (transformers, nemo) |
Transformers; NeMo | a few GB or CPU | Whisper large-v3 |
| Speech out | text-to-speech | model's own | model's own | a few GB or CPU | Kokoro-82M |
| Vision | image-classification, mask-generation, image-text-to-text | safetensors (transformers) |
Transformers; the language runtimes for VLMs | 87M-param models on CPU; VLMs as language models | SAM 3; Qwen3-VL |
| Embeddings | sentence-similarity, feature-extraction, text-ranking | safetensors (sentence-transformers) |
sentence-transformers; vLLM; Ollama | near zero | all-MiniLM-L6-v2; Qwen3-Embedding |
| Classical | none; not on the hub | pickle, ONNX, the library's own | scikit-learn and relatives | CPU | the scikit-learn user guide |
3.10The question Nexus leaves you with
You can choose a model for a job. What you cannot yet read is the release post that announced it: the training words, the architecture words, the words on the invoice, and the things bolted on around it. That is the Apex tier.