Tier 04 · Advanced

The Apex

For the person who reads release posts and vendor decks and wants to stop needing a dictionary.

4.1

How was it made, and what do the training words mean?

The DeepSeek-R1 card describes a model called R1-Zero trained by large-scale reinforcement learning with no supervised fine-tuning first, in which reasoning behaviours emerged on their own; then R1 proper, which adds a small set of cold-start examples before the reinforcement learning; then six smaller models distilled from R1 onto Llama Meta and Qwen bases. One card, four training words, and the whole 2025 story of reasoning models in it.

Supervised, unsupervised, self-supervised. Three ways of showing a model data. Supervised means each example carries a label; unsupervised means none do; the scikit-learn guide is organised by that split. Language-model pre-training is the third kind, self-supervised: the label is the next token, which the data supplies for free. That is why it can use internet-scale text corpora as training data and why nobody had to annotate it. Compute is the arithmetic spent on that training, usually counted in operations or GPU-hours. The scaling laws are the observed regularities between compute, data, parameter count and how wrong the model still is; release posts invoke them when they say a model is bigger.

Pre-training is the expensive part and produces the base model from the Nexus tier. It is not always done once: continued pre-training trains the base further on more or different data, for a domain or for longer context (the DeepSeek V4.1 card describes its context being extended late in pre-training), and it is still pre-training because the objective is the same. Everything aimed at behaviour after that is post-training, and the words below are its methods.

SFT, supervised fine-tuning: show the model examples of the behaviour wanted and nudge towards them. This is what turns a base model into an instruct model. Synthetic data, examples generated by other models, can appear at any stage of training and is especially common in post-training sets; the Phi-4 card describes its training blend as synthetic "textbook-like" data alongside filtered web text and acquired books.

RLHF, reinforcement learning from human feedback: people rank the model's answers, a reward model learns the ranking, and the model is trained to score well on it. The 2022 InstructGPT paper is the reference. RLAIF replaces the human ranking with feedback from a model against written principles; the 2022 Constitutional AI paper is the reference, and it is a family of methods rather than one swap. DPO, direct preference optimisation, from 2023, skips the reward model and trains directly on pairs of preferred and rejected answers; it is cheaper and, strictly, not reinforcement learning at all. GRPO, from the 2024 DeepSeekMath paper, is the reinforcement-learning algorithm DeepSeek used for R1, and RLVR, reinforcement learning with verifiable rewards, is the name for the recipe of rewarding answers that can be checked (maths, code, tests) rather than answers a person preferred. R1-Zero is the demonstration that, in DeepSeek's setup, this alone was enough for reasoning behaviours to emerge.

RLCD, reinforcement learning for calibrated decisions, is TypeSafe's name for the post-training behind Jev, the decision model in the box at 4.5, and it is here because the guide names it, marked reported like everything else about Jev. Its stated target differs from the two above: not a response people preferred, as in RLHF, and not a reward for verifiable correctness, as in RLVR, but a typed decision whose probabilities are optimised against outcomes, so that across many predictions the answers given 0.8 come true about 80 per cent of the time. That is a property of a population of predictions, not a promise about any one of them. TypeSafe describes the objective in its documentation; no weights and no paper detailed enough for reproduction have been published, so the mechanism and its results stay vendor-reported, as the box says.

Distillation: train a smaller student on a larger teacher's outputs. R1's card ships six such students and reports the 32B one outperforming a closed model of the time. The Qwen3.8-27B card, from the Origin tier, describes itself as bringing a Qwen-Max-class model to a compact dense form, which is the same move at a different scale.

Same word, different thing: distilled. Here it is a student model: a smaller one trained on a larger one's outputs. In image generation at 3.3 a distilled checkpoint is the same model retrained to need far fewer denoising steps. Both are a big thing made cheaper, and neither is the other. On a card, look at whether the parameter count changed.

Pruning: cut parameters out of a finished model. NVIDIA's Minitron-8B card describes pruning Nemotron-4 15B by embedding size, attention heads and MLP width, then distilling to recover. Pruned and distilled checkpoints look alike on a card and were made differently.

Quantisation-aware training, QAT: train with the quantisation in the loop so the low-bit checkpoint loses less. Gemma 4 Google DeepMind ships QAT variants as separate repositories with qat-q4_0 in the name, and Google publishes a launch post for them; the Vector tier's warning that quantisation is usually done by others has exceptions, and this is the clearest.

Alignment is the broad word for making a model's behaviour match what its makers intend, and post-training aimed at behaviour is the main tool for it. Guardrails are checks outside the model, before the input or after the output; they are a product feature, not a training method, and a card that mentions them is describing a system, not the file.

New here: training. How the nudging works, what a loss is, and why any of this converges belongs to neural networks. Nothing here depends on it.

One card, four training words

  1. R1-Zerolarge-scale reinforcement learning with no supervised fine-tuning first
  2. R1adds a small set of cold-start examples before the reinforcement learning
  3. six smaller modelsdistilled from R1 onto Llama and Qwen bases

Reasoning behaviours emerged on their own in the first, which is DeepSeek's claim for DeepSeek's setup.

When it happens

Pre-trainingis the expensive part and produces the base modelself-supervised: the label is the next token, which the data supplies for free
continued pre-trainingtrains the base further on more or different data, for a domain or for longer contextstill pre-training because the objective is the same
post-trainingEverything aimed at behaviour after that

synthetic data examples generated by other models, at any stage of training, and especially common in post-training sets

The base model the first stage produces is the one the Nexus tier is about.

How post-training is done: the methods

  • SFTshow the model examples of the behaviour wanted and nudge towards themturns a base model into an instruct model
  • RLHF2022people rank the model's answers, a reward model learns the ranking
  • RLAIF2022replaces the human ranking with feedback from a model against written principlesa family of methods rather than one swap
  • DPO2023skips the reward model and trains directly on pairs of preferred and rejected answersstrictly, not reinforcement learning at all
  • GRPO2024the reinforcement-learning algorithm DeepSeek used for R1
  • RLVRrewarding answers that can be checkedmaths, code, tests
  • RLCDa typed decision whose probabilities are optimised against outcomesvendor-reported

Done to a finished model, not aimed at behaviour

  • Distillationtrain a smaller student on a larger teacher's outputs
  • Pruningcut parameters out of a finished model
  • Quantisation-aware trainingtrain with the quantisation in the loop so the low-bit checkpoint loses less

Pruned and distilled checkpoints look alike on a card and were made differently.

Not a training method at all

  • Alignmentmaking a model's behaviour match what its makers intendthe goal, not a method
  • Guardrailschecks outside the model, before the input or after the outputa product feature, not a training method
Every word on a release post is one of these three: when it happened, how that stage was done, or not training at all.
4.2

What do the architecture words on a release post mean?

The Qwen3.8-27B card describes its layout as sixteen blocks of three Gated DeltaNet layers followed by one Gated Attention layer, with multi-token prediction trained in. The config lists 48 linear_attention layers and 16 full_attention. You do not choose any of these mechanisms when you use a checkpoint, but they decide which runtimes and kernels can run it, how much memory it needs and how fast it goes; and all of it appears on the release post.

One line each. Names only; the point is placement, not mechanism.

  • Attention family. Multi-head attention is the original. GQA (grouped-query attention) shares key and value heads between query heads; the Qwen config's 24 query heads to 4 key-value heads is it. MLA (multi-head latent attention) compresses the keys and values into a smaller latent. Sliding-window attention looks back a fixed distance. Linear attention and the DeltaNet variants replace the lookback and its KV cache with a fixed-size recurrent state, and hybrid stacks mix them with full attention, as Qwen's 3:1 layout does. Every one of these decides how much state the context costs, which is why they appear on a card at all.
  • MoE family. Routed experts (the Nexus tables' expert counts), top-k routing (the per-token counts), shared experts that every token uses (Kimi K3's card lists two), fine-grained experts (many small ones; Kimi's 896), and latent MoE, where the experts run in a compressed space, which a 2026 NVIDIA paper describes and Nemotron 3 uses; Kimi K3's own card names its version "Stable LatentMoE".
  • Sequence family. Transformer is the default: the layout of layers, from 2017, in which every position can look at every other, and nearly every model in this guide is one or a hybrid of one. Mamba and the state-space models are the recurrent alternative; hybrid Mamba-Transformer stacks put both in one model. Mixture-of-Transformers, on the Cosmos 3 card, is two transformer towers, one autoregressive and one diffusion, sharing a model.
  • Output tricks. Multi-token prediction (MTP) trains the model to predict several tokens ahead; its use at run time, speculative decoding, is a runtime feature and sits at 3.2.
  • Context tricks. RoPE (rotary position embedding, listed on the Qwen card with its dimension) is the position encoding on most current cards; context extension is any of several techniques (position scaling or interpolation, attention changes, further long-context training) for supporting longer inputs than the model was first trained on, which is why the Qwen card can say 262,144 natively and up to 1,000,000 extended, and why the two numbers differ.
  • Generation family. Autoregressive (one token at a time), diffusion (denoising in steps), DiT (a transformer doing the denoising), flow matching and rectified flow (related continuous generative formulations alongside diffusion, on the FLUX.2 and TRELLIS.2 cards), and non-autoregressive (a broad category of models that do not produce output one piece at a time; the decision model in the box at 4.5 is one example).

New here: architecture. What any of these do inside the arithmetic, the forward pass, belongs to neural networks. This list exists so the words have a shelf to sit on.

One block, repeated 16 times

Three linear layers, then one full one, the 3:1 layout.

The whole stack

16 blocks × 3 = 48 linear_attention, and 16 × 1 = 16 full_attention, which is exactly what the config lists.

Why the mix is there

  • Gated DeltaNetlinear attention: replaces the lookback and its KV cache with a fixed-size recurrent state
  • Gated Attentionfull attention: keeps the lookback and its KV cache, the state that grows with the context

Every one of these decides how much state the context costs, which is why they appear on a card at all.

And the heads on the same card

24 query:4 key-valuegrouped-query attention: 6 query heads share each key-value head
You do not choose any of this when you use a checkpoint. It decides which runtimes and kernels can run it, how much memory it needs and how fast it goes, and all of it appears on the release post. The point here is placement, not mechanism.
4.3

What did someone do to the base?

On the day this guide was checked, the hub held at least six repositories named some variant of Qwen3.8-27B-abliterated, from four publishers, one of them tagged uncensored. Every one I checked happened to keep Apache 2.0; the base licence permits that and does not require it, since Apache's clause 4 lets a publisher put additional or different terms on their own modifications while the original conditions still apply.

A fine-tune is the base trained further on someone's examples; a merge is two or more checkpoints averaged or spliced; a distilled checkpoint is a student of a larger model; an abliterated checkpoint has had its refusal behaviour removed by editing the weights directly; "uncensored" is a broader marketing word that may mean abliteration, a fine-tune, a dataset choice or something else, and the card is the only place that says which; a long-context variant has had its context stretched. All of them are files in the same format as the base, and the card's base_model field, when the publisher fills it in, is the provenance.

Licence inheritance is the thing to read. A permissive base (Apache 2.0, MIT) lets a derivative publisher choose terms for their additions, including keeping the same licence, while the base's own conditions carry through; a custom base licence (Llama's, Gemma 3's) binds the derivative to its terms, which is why the abliterated Llama 3.1 in the sources carries llama3.1 as its licence field and the abliterated Qwen Alibaba carries Apache. The publisher of a derivative is not the maker of the model, and their card is the only description of what they changed.

One base, five things someone can have done to it

the baseall of them are files in the same format as it
  • fine-tunethe base trained further on someone's examples
  • mergetwo or more checkpoints averaged or spliced
  • distilleda student of a larger model
  • abliteratedhad its refusal behaviour removed by editing the weights directly
  • long-contexthad its context stretched

The card's base_model field is the provenance, when the publisher fills it in.

And one word that is not a sixth

uncensoreda broader marketing word: it may meanabliterationa fine-tunea dataset choicesomething elsethe card is the only place that says which

What the derivative's licence can be

A permissive baseApache 2.0, MITlets a derivative publisher choose terms for their additions, including keeping the same licencewhile the base's own conditions carry throughthe abliterated Qwen carries Apache
a custom base licenceLlama's, Gemma 3'sbinds the derivative to its termsthe abliterated Llama 3.1 in the sources carries llama3.1 as its licence field

The publisher of a derivative is not the maker of the model, and their card is the only description of what they changed.

What one search turned up

  • at least sixrepositories named some variant of Qwen3.8-27B-abliterated
  • from fourpublishers
  • one of themtagged uncensored

Every one checked happened to keep Apache 2.0; the base licence permits that and does not require it.

Five different operations, one file format, and only the card to tell you which of them happened.
4.4

What is a world model, and what is a robot model?

NVIDIA's Cosmos 3 card, updated on 16 September 2026, calls itself an omnimodal world model for physical AI, built on a Mixture-of-Transformers architecture with an autoregressive tower and a diffusion tower, in Edge, Nano and Super sizes from 3.9 to 64.6 billion parameters under NVIDIA's OpenMDW licence, with a Nano variant published as a robot policy for the DROID dataset, per its repository name. Meta's V-JEPA 2 cards, from FAIR, describe a video model trained to predict in representation space. Physical Intelligence's π0.5, published through LeRobot, calls itself a vision-language-action model that executes long-horizon tasks in unseen environments.

World model. A model of an environment's state and how it changes, which may be conditioned on actions and may be trained without them (V-JEPA 2's cards describe video-only pre-training), used to predict what happens next: in pixels (generate the next frames given an action) or in latent space (predict the next representation, not the next pixel). The second is the JEPA idea, and the V-JEPA line is its reference; its cards describe extending that predictive pre-training objective to video. The first is the generative-video lineage, which Cosmos 3 sits in with its diffusion tower. A third lineage trains an agent inside the world model's simulation rather than the real environment, model-based reinforcement learning, which is a different use of reinforcement learning from the post-training methods at 4.1. Which lineage a paper belongs to is the first thing to work out, because the word is the same for all three.

VLA, vision-language-action: a model that takes images and instructions and outputs robot actions. π0 and π0.5 on the hub carry the robotics task tag and 3.5 to 3.6 billion parameters.

3D. Text-to-3D and image-to-3D produce a 3D representation, which may be a mesh, a Gaussian splat, a radiance field or voxels depending on the model; the task tag is image-to-3d. TRELLIS.2 (Microsoft, MIT licence, 4B, a flow-matching model over a sparse voxel structure per its card) and Hunyuan3D 2.1 (Tencent community licence) are the open anchors; SAM 3D Objects and SAM 3D Body are Meta's, under a custom licence with manual gating. NeRF and Gaussian splatting are two of those representations, not generators, and appear as the format such models output or consume.

What people get wrong. Treating a world model as a video generator with a joystick. Two of the three lineages never render a frame, and when generative video is used as a world model its purpose is predicting and simulating an environment's dynamics, not ordinary media generation.

One word, three lineages

A model of an environment's state and how it changes, which may be conditioned on actions and may be trained without them.

  • in pixelsgenerate the next frames given an actionthe generative-video lineageCosmos 3 sits in it, with its diffusion toweran autoregressive tower and a diffusion towerrenders frames
  • in latent spacepredict the next representation, not the next pixelthe JEPA ideathe V-JEPA line is its referencevideo-only pre-trainingnever renders a frame
  • trains an agent inside the world model's simulationrather than the real environmentmodel-based reinforcement learninga different use of reinforcement learningfrom the post-training methodsnever renders a frame

Two of the three lineages never render a frame. Which lineage a paper belongs to is the first thing to work out, because the word is the same for all three.

Which is why the usual shorthand is wrong

treating a world model as a video generator with a joystick:it describes one of the three, and even there, when generative video is used as a world model, its purpose is predicting and simulating an environment's dynamics, not ordinary media generation.

And the robot model beside it

inimagesinstructions
outrobot actions
VLArobotics3.5 to 3.6 billion parameters

And one more pair that gets swapped

text-to-3D, image-to-3Dgenerators; the task tag is image-to-3d
the representations they producea mesha Gaussian splata radiance fieldvoxelsNeRF and Gaussian splatting are two of those representations, not generators; they appear as the format such models output or consume
Work out which lineage a paper is in before anything else: the word is the same for all three, and two of them never draw a picture.
4.5

What gets bolted on, and what is the ladder of "-engineering" words?

Anthropic donated the Model Context Protocol to the Linux Foundation on 9 December 2025, into a new Agentic AI Foundation co-founded with Block and OpenAI, and said at the time that more than 10,000 public MCP servers existed. The Linux Foundation had taken in Google's Agent2Agent protocol six months earlier. Two standards, one for connecting an AI application or agent to external tools and data and one for agents talking to each other, and both are now nobody's product.

Prompting. Prompt engineering is what the model is told; few-shot is showing it examples in the prompt; chain of thought is asking it to write its working, which reasoning models now do unasked. Context engineering is the successor word, in wide use since 2025: what goes into the window, in what order, from where, and the prompt-caching rules at 4.6 are one reason order matters.

RAG, retrieval-augmented generation: search first, then generate with the results in the context. The pieces are the embeddings and rerankers from 3.7, a vector database to hold the vectors, chunking (how documents are cut before embedding) and grounding (keeping the answer tied to the retrieved evidence; citations are one way to show the reader which evidence was used). RAG versus fine-tuning is the standard confusion: RAG changes what the model can see, fine-tuning changes how it behaves, and the question "should I fine-tune to teach it our documents" usually has RAG as its answer. GraphRAG retrieves over a graph of entities and relations instead of, or beside, chunks and vectors; a semantic layer or ontology is the enterprise-data word for telling a model what "revenue" means before it writes a query. Neither graph is the graph in graph engineering below.

Same word, different thing: grounding. Here it means keeping an answer tied to the evidence that was retrieved for it. In vision at 3.6 it means pointing at where in the image something is. One is about being right, the other about being precise, and the two turn up a paragraph apart in the same release post.

Tools. Tool use or function calling is the model emitting a structured request that the harness executes. MCP standardises how an AI application or harness connects to external tools and data, which is the protocol's own description of itself; the model sees the result. A2A standardises how agents reach each other. Both live under the Linux Foundation's Agentic AI Foundation, whose founding contributions also included the AGENTS.md convention for agent instructions.

Same word, different thing: memory. An agent's memory is state kept outside the model's weights and made available again on later turns, often by putting it, or a retrieved part of it, back into the context. A GPU's memory is the VRAM the file sits in. Only the second is hardware memory, the capacity that decides what the device running the file can hold.

Guard models. Small classifiers that sit before the input or after the output and flag unsafe content or injection attempts. Meta ships Llama Guard 4 and Prompt Guard 2 as separate checkpoints and its Llama 4 card tells developers to deploy them alongside the model; they are the guardrails of 4.1 as files rather than as a feature.

Agents. An agent is a model in a loop with tools and a goal. The harness is the machine around the model: the loop, the prompts, the memory, the permissions, the sandbox; the model is swappable and the harness is the product. The 2026 vocabulary of harnesses, one line each: subagents (a harness spawning narrower harnesses), skills (packaged instructions and tools a harness can load), hooks (code that runs at fixed points in the loop), plugins, computer use and browser agents (a harness driving a screen), sandbox (where the harness is allowed to act), human-in-the-loop and approval gates, bounded autonomy and token budgets. Harness engineering is an increasingly used name for building that machine. A decision model, the box below, is what vendors now propose to put at the branch points inside it.

The coding agents named at 1.2 are harnesses of this kind, and often appear in three common shapes: a feature inside a code editor, a program in a terminal on your machine, and a service that takes a task and returns the finished change. The vocabulary above is the same in all of them; what differs is where the loop runs and who is watching it.

Prompt injection is the failure the permissions exist for: instructions that try to redirect the agent, hidden in something the model reads, such as a web page, an email or a tool's result. A model has no separate channel for "this is data, not a command"; it has only the window. It is not a bug in one product; it is what reading untrusted text with permission to act means, and the defences are all in the harness: which tools it may call, what it must ask before doing, where it is allowed to act, and a person watching. Tools are not required for prompt injection to matter: if injected content changes a summary, recommendation or decision that a person then acts on, the model has still carried the attack across the trust boundary. 4.7 places the word beside the other things that go wrong.

Emerging: a model that returns a decision, not text. On 15 September 2026 a company called TypeSafe AI announced a model named Jev that returns no text: it takes a block of program state and a set of typed questions, and returns one answer per question with a probability attached. TypeSafe calls the category a System One model, after the fast mode of thought in Daniel Kahneman's account, and names its training method RLCD, reinforcement learning for calibrated decisions, which the Apex tier's training section places at 4.1 beside the other post-training methods. Its documentation gives the questions three shapes: Choice, one option from a list; Score, a position on an ordered rubric; and Noul, whether a statement is true, as a probability from zero to one. Choice and Score return the probabilities across their options and a separate confidence figure, which summarises how concentrated those probabilities are; a Noul answer carries no confidence field. Questions asked against one state are answered independently of each other, and the documentation's own advice is to ask one narrow thing per question and combine the answers with logic in your code: the model supplies the judgement and the program keeps the control flow. Everything about it is reported. No weights and no paper detailed enough for reproduction have been published, its speed claims are its own measurements, and the independent reading is that it can still choose the wrong option and that giving up step-by-step reasoning caps what it can do. What is durable is the shape: a choice, a score or a yes-or-no, typed, with a probability attached, and nothing to parse. The word to hold on to is calibration: a calibrated probability is one where the events the model marks as 90 per cent likely come true about nine times in ten, across many predictions. That is a property of the probabilities, not of the confidence figure, which only says how peaked the distribution was; a model can emit numbers that look like probabilities without their being calibrated, and a vendor's claim of calibration is a claim to be tested on your own data. "Cannot hallucinate" is a claim about the shape of the output and says nothing about its correctness. Where such a thing would sit is here, at the branch points of a harness: which tool to call next, whether to retry, whether to escalate.

Emerging vocabulary. The rest of this section, the ladder and vibe coding, is current industry language rather than settled terminology; RAG, tool calling, MCP and A2A above are established. On the site this part is set apart visually.

The ladder. This is an emerging 2026 vocabulary, not a settled hierarchy, and every source for it is a blog or a vendor deck. On those decks the "-engineering" words arrive in this order, each claiming to sit above the last: prompt engineering → context engineering → harness engineering → loop engineering (one context, one think-act-observe cycle, the model deciding what comes next, with verification, persisted state and stopping rules around it) → graph engineering (the process made explicit as nodes, typed state and edges you can branch, parallelise and retry, across several agents or steps). The loop-versus-graph argument is the live one: most tasks are a well-scoped loop, and a graph earns its place when work splits into specialties, needs fan-out and fan-in, wants different models per step, or needs auditable control flow. Harness, loop and graph are three layers that stack rather than compete. On a deck, "agentic" means only that one of the above is present, and the word is an adjective. A workflow is the same loop with the steps fixed in advance by a person; an agent is one where the model chooses the next step.

Vibe coding is cultural rather than technical vocabulary and is here only because it is everywhere: it meant, in 2025, accepting generated code without reading it, and in 2026 usage has shifted towards specifications in, code out, with human review between. It names a practice rather than a model.

Test-time compute is the umbrella term for extra work done at answer time to get a better answer: reasoning tokens, sampling several candidates and picking one, search, verifier loops, tool calls, iterative refinement. A thinking budget is one implementation of it, the thinking-model dial from the Vector tier seen from the vendor's side: more tokens of working, more cost, sometimes better answers. Gemma 4's fixed budgets and gpt-oss's OpenAI low-medium-high are two ways of exposing that one.

This section treats RAG, tools and agents as language-system concepts, which is where the vocabulary comes from. Multimodal agents also retrieve images and call image, audio and video models as tools; the other families' own bolt-ons, adapters and controllers, were covered in the Nexus tier.

Search first, then generate

before a question is ever asked
  1. documents
  2. chunkinghow documents are cut before embedding
  3. embeddingsfrom § 3.7
  4. vector databaseholds the vectors
and then, per question
  1. the question
  2. retrievesearch first
  3. rerankersfrom § 3.7
  4. generatewith the results in the context

Grounding keeps the answer tied to the retrieved evidence. Citations are one way to show the reader which evidence was used.

The standard confusion

RAGchanges what the model can see
fine-tuningchanges how it behaves

"Should I fine-tune to teach it our documents?" usually has RAG as its answer.

Two neighbours

  • GraphRAGretrieves over a graph of entities and relations instead of, or beside, chunks and vectors
  • semantic layeror ontology, the enterprise-data word for telling a model what "revenue" means before it writes a query

Neither graph is the graph in graph engineering, which is the different thing in the figure below.

How the decks order the words

  1. prompt engineeringwhat the model is told
  2. context engineeringwhat goes into the window, in what order, from where
  3. harness engineeringbuilding the machine around the model
  4. loop engineeringone context, one think-act-observe cycle, the model deciding what comes next, with verification, persisted state and stopping rules around it
  5. graph engineeringthe process made explicit as nodes, typed state and edges you can branch, parallelise and retry, across several agents or steps

Each claiming to sit above the last.

The live argument

loopmost tasks are a well-scoped loop
graphearns its place when
  • work splits into specialties
  • it needs fan-out and fan-in
  • it wants different models per step
  • it needs auditable control flow

And the correction that goes with it

harnessloopgraph

Three layers that stack rather than compete.

4.6

What are the words on the invoice?

OpenAI's prompt-caching documentation, as read on the day this guide was checked, says that a cache hit needs the whole rendered prefix to match, that a prefix must reach a minimum length before it can be cached, that cached reads are billed at a fraction of the uncached input rate while on some models a cache write costs more than an uncached read, and that on newer models developers can place cache breakpoints themselves. The minimums, ratios and controls are per model generation and change; the page is the source and this guide does not repeat its numbers. What does not change is the consequence: the order of a prompt is now a cost decision.

No prices appear in this guide. Ratios and rules do.

  • Input versus output. Two prices, and output is usually the dearer; pricing is the provider's decision, but generating a token costs more than reading one (Vector), and prices tend to follow. Reasoning tokens are output; a thinking model's bill is mostly its working.

Same word, different thing: cache. Prompt caching is a provider reusing the computed prefix of a prompt to bill and respond faster. The KV cache is the runtime's working memory for the context, sitting in VRAM. One is a discount, the other is a cost.

  • Prompt caching. Provider-side. The repeated prefix (instructions, tool definitions, documents, conversation so far) is billed at a fraction on later calls if it matches exactly and sits at the front. Put the stable material first and the variable material last; the provider's response reports a cached-token count so you can check whether it worked.
  • Semantic caching. Application-side. Similar questions served from a stored answer, at a similarity threshold you choose, with the risk that "similar" is wrong.
  • Batch. A file of requests submitted together and returned within a window measured in hours rather than seconds, at a reduced rate; OpenAI's batch endpoint is the reference shape, and its limits are on its page. Anything that can wait belongs there.
  • Structured outputs. The model constrained to a schema, which most provider APIs offer either through function calling or a schema-constrained response format. It cuts parsing and schema failures; it does not by itself cut output tokens. It is not a decision model, because the tokens are still generated one at a time and the schema constrains the shape, not the reasoning.
  • Rate limits. Per-minute ceilings on requests and tokens, set per account; the number that decides whether your traffic fits before the price does.
  • Gateway and router. One endpoint in front of many providers. OpenRouter is the public example and publishes which models support which parameters; LiteLLM and the cloud vendors' gateways are the self-hosted and enterprise ones. Model routing, sending each request to the cheapest model that can answer it, is the cost lever that gateways make possible.
  • Streaming. Billing-neutral, latency-visible; the choice between a spinner and a first token.
  • Self-hosting versus API. GPU hours against tokens. The crossover depends on utilisation: a card that is busy all day is cheap per token and a card that is idle is the most expensive way to run a model. The ratios are yours to compute; the guide gives no price and no crossover figure because both change monthly.

One prompt, ordered for the cache

stable: the repeated prefix
instructionstool definitionsdocumentsconversation so far
variable
this turn

Put the stable material first and the variable material last.

What a hit needs

  • the whole rendered prefix has to match
  • it has to sit at the front
  • a prefix must reach a minimum length before it can be cached
  • on newer models, developers can place cache breakpoints themselves

Cached reads are billed at a fraction of the uncached input rate, and on some models a cache write costs more than an uncached read. The minimums, ratios and controls are per model generation and change, so the provider's own page is the source for them, not this guide.

The provider's response reports a cached-token count, so you can check whether it worked.

Same word, different thing

  • prompt cachea provider reusing the computed prefix of a prompt, to bill and respond fastera discount
  • KV cachethe runtime's working memory for the context, sitting in VRAMa cost
The order of a prompt is now a cost decision.
4.7

How is it evaluated, and what goes wrong?

The MTEB paper that defined the embedding benchmark reported, on its own results, that no single method dominated across tasks. That sentence generalises to every leaderboard in this guide.

Benchmarks and arenas. A benchmark is a fixed test set with a score; an arena is people voting between two anonymous outputs, producing an Elo-style ranking, a rating of the kind chess uses. MTEB for embeddings, the vendor tables on every language-model card, arenas for images and video. Benchmark contamination is the test set having leaked into the training data, and it is one reason a leaderboard score does not survive contact with your own data; the other is that the benchmark measured something other than your task. An eval harness is the program that runs a benchmark against a model; two harnesses give two scores for one checkpoint, because settings differ. EleutherAI's lm-evaluation-harness is the reference example: a runner that applies the same tasks and settings to different models, and the backend of Hugging Face's Open LLM Leaderboard. An eval, as a noun, is the industry's short word for a benchmark run, and for the discipline of running them. LLM-as-judge is scoring outputs with another model, which is cheap and inherits the judge's biases.

Same word, different thing: harness. At 4.5 a harness is the machine around an agent: the loop, the tools, the permissions. Here it is the program that runs a benchmark. Both are the scaffolding around a model, and neither is the model.

Model card versus system card. The card describes the file; a system card describes what the vendor tested a deployed system for. gpt-oss's card, for one, is explicit that the chain of thought is for debugging and not for end users, which is a system statement on a model card.

What goes wrong. Hallucination (Origin) and grounding, the practice of tying an answer to retrieved or cited material, which reduces it without being its opposite. Prompt injection: untrusted instructions embedded in content the model reads, attempting to redirect the model or the agent; in an agent harness the content can be a tool's description or result, which is the form to watch for. Jailbreak: an input crafted to bypass a model's or system's behavioural and safety constraints. Red-teaming is the deliberate search for such inputs before release, and a system card usually reports it. Sycophancy is a model agreeing with what the user appears to want; it is a post-training effect, not a fact about the file's knowledge. Regression, drift and decay are three words for a model getting worse, and published glossaries do not agree on which is which; one usable split is regression for a new checkpoint doing worse on something the old one did, drift for behaviour changing under a fixed checkpoint because the inputs changed, and decay for a vendor's hosted model changing under you. State your own definitions before you use them in a report.

Provenance of output. A deepfake is generated or edited media presented as real, and the provenance mechanisms exist because of it. Two mechanisms, and they fail differently. C2PA Content Credentials are cryptographically signed metadata recording a file's origin and edits, an open specification from the Coalition for Content Provenance and Authenticity; Google's cloud documentation describes it and states that a modification by a non-C2PA tool breaks validation. A watermark is a signal embedded in the pixels or audio themselves; Google DeepMind's SynthID is the reference; DeepMind's own page describes it as embedding an imperceptible watermark into AI-generated images, audio, text or video and scanning for it, and calls it a beta toolkit and not a silver bullet. A screenshot normally loses the original C2PA metadata; a robust watermark may survive transformations such as screenshots or recompression, depending on the scheme, but needs the vendor's detector to be read. Detector products that work from the content alone have neither a signature nor a watermark to check; whatever they report is a probabilistic inference, and the two mechanisms above are the stronger evidence, not the only kind. The reason the words exist in 2026 is law: Article 50 of the EU AI Act, whose marking and labelling obligations apply from 2 August 2026, with a voluntary Code of Practice on Transparency of AI-Generated Content published by the European Commission on 10 June 2026; both dates are from the Commission's own pages. The Act's text in the Official Journal is the source for anything beyond the dates.

Where a score comes from

  1. A benchmarka fixed test set with a score
  2. An eval harnessthe program that runs a benchmark against a model
  3. its settingstwo harnesses give two scores for one checkpoint, because settings differ
  4. the scoreone number, from that harness, on that test set

Change the harness or its settings and the number changes, with the checkpoint untouched.

an arenapeople voting between two anonymous outputs, producing an Elo-style ranking
LLM-as-judgescoring outputs with another modelcheap and inherits the judge's biases

Two reasons the number does not survive contact with your data

the training data
the test set
Benchmark contamination: the test set having leaked into the training data
the benchmark was not your taskthe benchmark measured something other than your task

Three words for a model getting worse

  • regressiona new checkpoint doing worse on something the old one did
  • driftbehaviour changing under a fixed checkpoint because the inputs changed
  • decaya vendor's hosted model changing under you

Published glossaries do not agree on which is which. This is one usable split, not a definition. State your own definitions before you use them in a report.

No single method dominated across tasks. That is the MTEB paper's finding on its own results, and it generalises to every leaderboard in the guide.

Two mechanisms, and they fail differently

alongside the fileC2PA Content Credentialscryptographically signed metadata recording a file's origin and editsan open specification from the Coalition for Content Provenance and Authenticity
inside the contentA watermarka signal embedded in the pixels or audio themselvesGoogle DeepMind's SynthID is the reference: embedding an imperceptible watermark into AI-generated images, audio, text or video and scanning for it

Which is the whole of it: one rides beside the file, the other is in the file. Everything below follows from that.

So they break differently

  • any non-C2PA tool edits ita modification by a non-C2PA tool breaks validation
  • someone takes a screenshotA screenshot normally loses the original C2PA metadata
  • a screenshot, or recompressiona robust watermark may survive transformations such as screenshots or recompression, depending on the scheme
  • and even thenit needs the vendor's detector to be read

SynthID's own page calls it a beta toolkit and not a silver bullet.

And the third thing, which is not a third mechanism

Detector products that work from the content alonehave neither a signature nor a watermark to check; whatever they report is a probabilistic inference.The two mechanisms above are the stronger evidence, not the only kind.

Why the words exist at all

  • 2 August 2026marking and labelling obligations apply, Article 50 of the EU AI Act
  • 10 June 2026a voluntary Code of Practice on Transparency of AI-Generated Content published by the European Commission

Both dates are from the Commission's own pages. The Act's text in the Official Journal is the source for anything beyond the dates.

A screenshot is the test: it takes the metadata off and may leave the watermark on.
4.8

How do you read the table?

A release post's table has a shape, and the shape is read the way the card was read at 2.13: in an order, and with the vendor's claim last. First the columns: who is in the comparison class, and who is not. A post compares against the models the maker chose; a missing rival is information, and so is a rival's older version standing in for its current one. Then the rows: benchmark names, and a reader will keep meeting the same dozen; each is a fixed test set with a score, and its name tells you the kind of task (knowledge questions, maths competitions, code repairs, agentic tool use) and nothing about your task. Then the footnotes, which are where the settings live: "pass@1" means one attempt was scored; "n-shot" means the model was shown n worked examples first; "with tools" means it could run code or search; "high" or "max" reasoning effort means the most working it is allowed; "internal eval" means the maker wrote the test and nobody else has run it. Two harnesses give two scores for one checkpoint, and a vendor's number was produced on the vendor's harness with the vendor's settings. A bold cell is the vendor's emphasis, not a measurement. A blank cell is a decision. None of this makes the table worthless; it makes it a claim, which is what the guide has called it since the intro.

Independent leaderboards exist, run the same tests on settings that are theirs, and are the second thing to check; the arena at 4.7 is one kind. The names a reader keeps meeting, with the kind of task each is: MMLU, multiple-choice questions across 57 subjects from elementary mathematics to law; MMLU-Pro, the same idea with the trivial questions removed and ten options instead of four; GPQA, 448 questions written by domain experts in biology, physics and chemistry; SWE-bench, a codebase and a description of an issue, where the model must edit the code to resolve it; Humanity's Last Exam, 2,500 questions across dozens of subjects written by subject-matter experts; and an arena such as Chatbot Arena, people voting between two anonymous outputs, ranked from the pairwise votes. A multiple-choice score, a code repair and a vote are three different kinds of number, and a table that puts them in one column is asking you not to notice.

The shape of the table, with nothing in it

this modela rivala rival's older versionnot compared
MMLUscorescorescore
MMLU-Proscore†score
GPQAscorescore
SWE-benchscore‡§scorescore
Humanity's Last Examscore¶scorescore
  • †pass@1one attempt was scored
  • ‡n-shotshown n worked examples first
  • §with toolscould run code or search
  • internal evalthe maker wrote the test

Read it in this order

  1. Who is in the comparison class, and who is not.
  2. A missing rival is information.
  3. A bold cell is the vendor's emphasis, not a measurement.
  4. A blank cell is a decision.

Columns first, then rows, then the footnotes: the shape is read the way the card was read, and with the vendor's claim last.

No model is named and no number is real. The cells hold the kind of thing a cell holds; what a real table holds is a claim.
4.9

Whose names are these?

  • Open-weight labs, by the anchors in this guide: Alibaba (Qwen, Qwen-Image, Wan), DeepSeek, Moonshot (Kimi), Zhipu (GLM), Meta (Llama, SAM, V-JEPA), Google DeepMind (Gemma, SigLIP), OpenAI (gpt-oss, Whisper, CLIP), Mistral, NVIDIA (Nemotron, Parakeet, Canary, Cosmos), Black Forest Labs (FLUX), Stability (Stable Diffusion, Stable Audio), Lightricks (LTX), Tencent (Hunyuan), Microsoft (Phi, Florence, TRELLIS), Physical Intelligence (π), Kyutai, Sesame, Nari.
  • Closed labs, named as product makers at 1.2 and as standards donors at 4.5, and otherwise absent by scope: OpenAI, Anthropic, Google DeepMind.
  • Hubs and runtimes. Hugging Face and ModelScope; Transformers, Diffusers, sentence-transformers, NeMo; llama.cpp, Ollama, LM Studio; vLLM, SGLang, TensorRT-LLM; ComfyUI.
  • Frameworks. PyTorch is what almost every card in this guide lists; JAX and TensorFlow are the others. They are general model-development frameworks used for training and inference; the specialised runtimes named elsewhere in this guide are usually what a reader meets at deployment.
  • Agent and orchestration frameworks. LangChain and LangGraph, Semantic Kernel, Google ADK, CrewAI, the vendors' Agents SDKs. Names only; 4.5 says why the model is absorbing part of their job.
  • Standards bodies. The Open Source Initiative (the definition at 2.11), the Linux Foundation's Agentic AI Foundation (MCP, A2A, AGENTS.md), the C2PA.

A large share of the open-weight models in this guide's tables are from Chinese labs. That is a fact about 2026's release schedule, and it says nothing about the licence on any one of them; read the card.

4.10

The families, final column

The families again, with two more things about each: how the family's models are trained, and what gets bolted on.

Family Anchor Trained by Bolted on
Language Qwen3.8-27B pre-training, SFT, RL (RLHF, RLVR), distillation RAG, tools, MCP, agents, harnesses, gateways
Image FLUX.2 klein 4B diffusion or flow-matching training, distillation for turbo variants LoRA, ControlNet, upscalers, C2PA and watermarks
Video Wan 2.2 TI2V-5B as image, with MoE in some LoRA, offloading, C2PA and watermarks
Speech in Whisper large-v3 supervised on labelled audio VAD, diarisation, inverse text normalisation
Speech out Kokoro-82M supervised on paired text and audio voice cloning, vocoders
Vision SAM 3; Qwen3-VL contrastive (CLIP family), supervised, VLM post-training grounding, OCR pipelines
Embeddings all-MiniLM-L6-v2; Qwen3-Embedding contrastive on pairs vector databases, rerankers, hybrid search
Classical scikit-learn supervised or unsupervised on tables feature pipelines
World and robot Cosmos 3; V-JEPA 2; π0.5 video and representation prediction; policy training simulators, robot stacks
4.11

Where to go next

Every box marked "New here" stopped at the edge of one of these subjects. None of them is needed to read a filename, a card or a release post, which is why this guide stops where it does. They are named here so that you know what to look up.

  • Video memory, for what fit and bandwidth are and why they decide what a card can run
  • The graphics card, for the memory ladder and what a tensor core is
  • Attention and the KV cache, for how the next piece is computed and why the context costs memory
  • Neural networks, for how training moves the numbers and what the architecture words do inside the arithmetic
  • Tokenizers, for how the dictionary is built and why two models cut the same sentence differently
  • Quantisation as a craft, for how a quantiser chooses which parameters keep more bits and what calibration data is
  • And as systems rather than files: retrieval, agents and the protocols between them, the image pipeline, speech end to end, world models and physical AI