The Reference
Every figure in this guide was checked against the source listed for it on 20 September 2026. Where I could find only a vendor's own claim, the text says "reported". Where I could not find a source, the figure is not here.
The Index
Each entry names the tier and the section that places the term.
A
- A2A: Apex, What gets bolted on
- Abliterated: Apex, What did someone do to the base
- Active parameters: Vector, What does "235B-A22B" mean
- Adapter: Vector, What is a checkpoint
- Agent, harness, subagent, skill, hook, sandbox: Origin, Where do you meet one; Apex, What gets bolted on
- Agentic, workflow: Apex, What gets bolted on
- Alignment, guardrails: Apex, How was it made
- API: Origin, Where do you meet one; Vector, What is inference
- API-only: Vector, Are open weights open source
- Arena, benchmark, contamination: Apex, How is it evaluated
- Artificial intelligence, AI: Intro, What is "artificial intelligence"; Origin, What does "AI" actually mean on the box
- ASR, STT, WER: Nexus, Speech and audio
- Attention (GQA, MLA, sliding window, linear, DeltaNet): Apex, What do the architecture words on a release post mean
- AWQ, GPTQ, FP8, NVFP4, MXFP4: Vector, Why does one model have thirty downloads; Nexus, Language models: which variant
B
- Base, Instruct, Chat, Thinking: Nexus, Language models: which variant
- Batch (as a speed): Vector, What is inference; (as an API): Apex, What are the words on the invoice
- Benchmark table, pass@1, n-shot: Apex, How do you read the table
- BF16, FP16, INT8, INT4: Vector, Why does one model have thirty downloads
- Bit, byte, gigabyte: Origin, Why is it a download
- Borrowed words (learn, know, understand, think, reason, remember): Origin, How did it learn
- Bytes per parameter: Vector, Why does one model have thirty downloads
C
- C2PA, watermark, SynthID: Apex, How is it evaluated
- Calibration, confidence: Apex, What gets bolted on; How was it made
- Cap, plan: Origin, What is it actually reading
- CFG, guidance, sampler, seed, negative prompt: Nexus, Image generation
- Chain of thought, few-shot: Apex, What gets bolted on
- Chat product: Origin, Where do you meet one
- Chat template, special tokens: Vector, What is in the folder
- Chatbot: Origin, Where do you meet one
- Checkpoint: Vector, What is a checkpoint
- Choice, Score, Noul (TypeSafe): Apex, What gets bolted on
- Classifier, regressor, clustering, anomaly detection: Nexus, Classical machine learning
- CLIP, SigLIP, ViT, contrastive: Nexus, Vision understanding
- Cloud: Origin, Where do you meet one
- Coding agent: Origin, Where do you meet one; Apex, What gets bolted on
- ComfyUI: Nexus, Image generation
- Computer use, browser agent: Apex, What gets bolted on
- Confabulation: Origin, How did it learn
- Context engineering: Apex, What gets bolted on
- Context window, context length: Vector, What is a token; Nexus, Language models: which variant
- ControlNet: Nexus, Image generation
- Conversational speech model: Nexus, Speech and audio
- Cosine similarity, dimensions, Matryoshka (MRL): Nexus, Embeddings and rerankers
- CPU, processor: Origin, What does it run on
- CUDA: Origin, What does it run on
D
- Decision model, System One (emerging): Apex, What gets bolted on
- Decode, prefill, time to first token, tokens per second: Nexus, Language models: the dials
- Deep learning, neural network, layer: Origin, What does "AI" actually mean on the box
- Deepfake: Apex, How is it evaluated
- Dense: Vector, What does "235B-A22B" mean
- Diarisation, VAD: Nexus, Speech and audio
- Diffusion, DiT, flow matching, rectified flow: Nexus, Image generation; Apex, What do the architecture words on a release post mean
- Distillation: Apex, How was it made
- DPO, GRPO, RLVR, RLHF, RLAIF: Apex, How was it made
- Drift, regression, decay: Apex, How is it evaluated
E
- Effective parameters: Nexus, Language models: which variant
- Embedding (the layer): Vector, What is a parameter; (the model): Nexus, Embeddings and rerankers
- Eval, eval harness, LLM-as-judge: Apex, How is it evaluated
- Everyday AI, spam filter, translation: Origin, The AI you met today
F
- Fine-tuning: Vector, What is a checkpoint; Apex, What did someone do to the base
- Foundation model, frontier model: Vector, What does "multimodal" mean
- Front, front end: Origin, Where do you meet one; Vector, What is in the folder
- Function calling, tool use: Apex, What gets bolted on
G
- Gated download: Vector, What is in the folder
- Gateway, router, OpenRouter: Apex, What are the words on the invoice
- Generative, generative AI: Origin, What does "AI" actually mean on the box
- GGUF: Vector, What is in the folder
- GPT, generative pre-trained transformer: Origin, What does "AI" actually mean on the box
- GPU, graphics card: Origin, What does it run on
- Graph engineering, loop engineering: Apex, What gets bolted on
- GraphRAG, knowledge graph, semantic layer: Apex, What gets bolted on
- Grounding: Nexus, Vision understanding; Apex, How is it evaluated
- Guard models (Llama Guard, Prompt Guard): Apex, What gets bolted on
H
- Hallucination: Origin, How did it learn
- Harness (agent, eval): Apex, What gets bolted on; How is it evaluated
- Hosted: Vector, What is inference
- Hybrid search, semantic search: Nexus, Embeddings and rerankers
I
- Inference: Origin, What does it do when it runs; Vector, What is inference
- Inpainting, outpainting, image-to-image: Nexus, Image generation
- Input tokens, output tokens, thinking tokens, cached tokens: Vector, What is a token
J
- Jailbreak, red-teaming, sycophancy: Apex, How is it evaluated
- JAX, PyTorch, TensorFlow: Vector, What is in the folder; Apex, Whose names are these
- JEPA: Apex, What is a world model
K
- Kernel: Vector, Why does one model have thirty downloads
- Knowledge cutoff: Origin, How did it learn; Vector, What is in the folder
- KV cache: Nexus, Language models: which variant
L
- Latency, throughput: Vector, What is inference
- Latent MoE, shared experts, fine-grained experts: Apex, What do the architecture words on a release post mean
- Licence (Apache, MIT, CC, community, non-commercial): Vector, Are open weights open source; every Nexus licence trap
- LLM, large language model: Origin, What does "AI" actually mean on the box; What does it do when it runs; Vector, What is a parameter
- Local, on-device: Origin, Where do you meet one; Vector, What is inference
- Logits, probability distribution, greedy decoding, seed, determinism: Nexus, Language models: the dials
- Loop: Origin, What does it do when it runs
- LoRA: Vector, What is a checkpoint
M
- Machine learning, training data: Origin, What does "AI" actually mean on the box
- Mamba, state-space, Mixture-of-Transformers: Apex, What do the architecture words on a release post mean
- MCP: Apex, What gets bolted on
- Memory (an agent's, a product's): Origin, How did it learn; Apex, What gets bolted on
- Memory floor: every Nexus family section
- Merge: Apex, What did someone do to the base
- MMLU, MMLU-Pro, GPQA, SWE-bench: Apex, How do you read the table
- mmproj, projector: Vector, Why does one model have thirty downloads; What does "multimodal" mean
- Modality, multimodal, omni: Vector, What does "multimodal" mean
- Model (the plain sense): Intro, What is "artificial intelligence"; Origin, Why is it a download
- Model card, system card: Vector, What is in the folder; Apex, How is it evaluated
- MoE, experts, routing, top-k: Vector, What does "235B-A22B" mean
- MTEB: Nexus, Embeddings and rerankers
- MTP, multi-token prediction: Nexus, Language models: the dials; Apex, What do the architecture words on a release post mean
N
- Name fragments (date stamps, patch16, so400m, hiera, TDT, 12Hz, n-gram embedding, Engram, effective): Vector, What does
Qwen3.8-27B-UD-Q4_K_M.ggufactually say - Narrow AI, general AI, AGI: Intro, What is "artificial intelligence"; Vector, What does "multimodal" mean
- NPU: Origin, What does it run on
O
- OCR, document understanding: Nexus, Vision understanding
- Offloading: Origin, What does it run on; Nexus, Video generation
- Ollama, LM Studio, llama.cpp, vLLM, SGLang: Nexus, Language models: which variant
- Open weight, open source, OSAID: Vector, Are open weights open source
- OpenAI-compatible endpoint: Nexus, Language models: which variant
P
- Parameter, weight, tensor: Vector, What is a parameter
- Post-training, pre-training, SFT: Apex, How was it made
- Prompt, completion: Origin, What is it actually reading; Vector, What is inference
- Prompt caching, semantic caching: Apex, What are the words on the invoice
- Prompt injection: Apex, What gets bolted on; How is it evaluated
- Pruning: Apex, How was it made
Q
- QAT, quantisation-aware training: Apex, How was it made
- QLoRA: Vector, What is a checkpoint
- Quantisation, Q4_K_M and the tag zoo: Vector, Why does one model have thirty downloads
R
- RAG, chunking, vector database: Apex, What gets bolted on
- Rate limits: Apex, What are the words on the invoice
- Reasoning model, thinking budget, test-time compute: Vector, What is a thinking model; Nexus, Language models: which variant; Apex, What gets bolted on
- Reranker, bi-encoder, cross-encoder: Nexus, Embeddings and rerankers
- RLCD, reinforcement learning for calibrated decisions (reported): Apex, How was it made; What gets bolted on
- RoPE, context extension: Apex, What do the architecture words on a release post mean
- Runtime, engine, backend: Vector, What is inference
S
- safetensors: Vector, What is in the folder
- Segmentation, detection, classification: Nexus, Vision understanding
- Self-hosted, provider-hosted: Origin, Where do you meet one; Vector, What is inference
- Self-hosting versus API: Apex, What are the words on the invoice
- Serving: Vector, What is inference
- Small language model, SLM: Nexus, Language models: which variant
- Spectrogram: Nexus, Speech and audio
- Speculative decoding, draft model: Nexus, Language models: the dials
- Step: Nexus, Image generation (same word, different thing)
- Streaming: Apex, What are the words on the invoice
- Structured outputs, JSON mode: Apex, What are the words on the invoice
- Supervised, unsupervised, self-supervised: Apex, How was it made
- Synthetic data: Apex, How was it made
- System instructions, developer instructions: Vector, What is in the folder
T
- Temperature, top-p, min_p, repetition penalty: Nexus, Language models: the dials
- Text encoder, VAE: Nexus, Image generation
- Thinking model, reasoning model: Vector, What is a thinking model
- Token: Origin, What is it actually reading; Vector, What is a token
- Tokenizer, vocabulary, BPE: Vector, What is a token
- Training: Origin, How did it learn; Apex, How was it made
- Transformer: Apex, What do the architecture words on a release post mean
- TTS, voice cloning, vocoder: Nexus, Speech and audio
U
- Uncensored: Apex, What did someone do to the base
- Unified memory: Origin, What does it run on
- Upscaler: Nexus, Image generation
V
- VAE: Nexus, Image generation
- Vector: Vector, Which words mean two things; Nexus, Embeddings and rerankers
- Vibe coding: Apex, What gets bolted on
- VLA, robot policy: Apex, What is a world model
- VLM: Nexus, Vision understanding
- VRAM: Origin, What does it run on
W
- World model: Apex, What is a world model
Z
- Zero-shot: Nexus, Vision understanding
The Sources
IntroThe Intro
- J. McCarthy, M. L. Minsky, N. Rochester and C. E. Shannon, "A Proposal for the Dartmouth Summer Research Project on Artificial Intelligence", dated 31 August 1955, as hosted on John McCarthy's Stanford pages (the phrase's first appearance, and the conjecture the intro paraphrases): http://www-formal.stanford.edu/jmc/history/dartmouth/dartmouth.html
OriginTier 01
- Qwen3.8-27B repository (file listing,
config.json,model.safetensors.index.json,README.md,LICENSE), Hugging Face Hub, read via the Hub API: https://huggingface.co/Qwen/Qwen3.8-27B - Token counts in "What is it actually reading?" were produced by running the
tokenizer.jsonfrom that repository with the Hugging Facetokenizerslibrary on the sentences shown; the Urdu and Arabic sentences are my own renderings of the English one - OpenAI, "What are tokens and how to count them?" (rules of thumb; input, output, cached and reasoning token categories): https://help.openai.com/en/articles/4936856-what-are-tokens-and-how-to-count-them
VectorTier 02
- PyTorch Foundation, "PyTorch Foundation Announces Safetensors as Newest Contributed Project to Secure AI Model Execution" (Paris, 8 April 2026; safetensors joins the Foundation as a hosted project): https://pytorch.org/blog/pytorch-foundation-announces-safetensors-as-newest-contributed-project-to-secure-ai-model-execution/
- Qwen3.8-27B repository as above; tensor shapes read from the safetensors headers of shards 1 and 3; special tokens from
tokenizer_config.json; template length fromchat_template.jinja - Unsloth GGUF quantisations of Qwen3.8-27B (file list and sizes, licence, download count), read via the Hub API: https://huggingface.co/unsloth/Qwen3.8-27B-GGUF
- Qwen3.8-Flash-Next model card, benchmark table, read 24 September 2026 (Qwen3.8-27B at 27B against Qwen3.7-Plus at 397B: the 27B scores above it on DeepSWE 1.1, SWE-bench Pro, CoWorkBench, JobBench, IFBench and LiveCodeBench v6, and below it on SWE-bench Multilingual and GPQA Diamond): https://huggingface.co/Qwen/Qwen3.8-Flash-Next
- Fragments box at 2.1, read 24 September 2026: Qwen-Image-2512 card ("the December update"): https://huggingface.co/Qwen/Qwen-Image-2512
- Fragments box at 2.1, read 24 September 2026: Mistral's models overview (Mistral Small 4 listed as v26.03, API name
mistral-small-2603): https://docs.mistral.ai/getting-started/models/models_overview - Fragments box at 2.1, read 24 September 2026: DeepSeek-V4-Pro-0813 card (the suffix is not explained there): https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro-0813
- Fragments box at 2.1, read 24 September 2026: google/vit-base-patch16-224 card ("a sequence of fixed-size patches (resolution 16x16)"): https://huggingface.co/google/vit-base-patch16-224
- Fragments box at 2.1, read 24 September 2026: Alabdulmohsin et al., "Getting ViT in Shape" (SoViT-400m, the shape-optimised ViT): https://arxiv.org/abs/2305.13035
- Fragments box at 2.1, read 24 September 2026: facebookresearch/hiera ("Hiera: A Hierarchical Vision Transformer without the Bells-and-Whistles"): https://github.com/facebookresearch/hiera
- Fragments box at 2.1, read 24 September 2026: facebookresearch/sam2 (the
sam2.1_hieracheckpoints): https://github.com/facebookresearch/sam2 - Fragments box at 2.1, read 24 September 2026: Xu et al., "Efficient Sequence Transduction by Jointly Predicting Tokens and Durations" (the Token-and-Duration Transducer): https://arxiv.org/abs/2304.06795
- Fragments box at 2.1, read 24 September 2026: Qwen3-TTS-12Hz-1.7B-Base card (its table lists the 12Hz tokenizer at 12.5 frames per second): https://huggingface.co/Qwen/Qwen3-TTS-12Hz-1.7B-Base
- Fragments box at 2.1, read 24 September 2026: Qwen3.8-Flash-Next card (the n-gram embedding: "indexing with short n-grams", "more amenable to offloading than Mixture-of-Experts"): https://huggingface.co/Qwen/Qwen3.8-Flash-Next
- Fragments box at 2.1, read 24 September 2026: DeepSeek-V4.1-Flash card ("Engram conditional memory (196B parameters, sparsely accessed via token-based lookup)"; 8B active in prefill, 16B in decode): https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash
- Fragments box at 2.1, read 24 September 2026: gemma-4-E4B-it card ("4.5B effective (8B with embeddings)", per-layer embeddings): https://huggingface.co/google/gemma-4-E4B-it
- Qwen3.8-27B file listing via the Hub API, read 24 September 2026 (
crc32.txt; an LFSoidhash on every large file): https://huggingface.co/api/models/Qwen/Qwen3.8-27B/tree/main - Hugging Face Hub, model cards documentation and the model card metadata specification, read 24 September 2026 (no cutoff key among the metadata): https://huggingface.co/docs/hub/model-cards and https://raw.githubusercontent.com/huggingface/hub-docs/main/modelcard.md
- Claude API documentation, context windows, read 24 September 2026 ("If the input alone already exceeds the model's context window, the API returns a 400 invalid_request_error"): https://platform.claude.com/docs/en/build-with-claude/context-windows
- llama.cpp server README, read 24 September 2026 (a
truncatedflag when prompt plus generated tokens exceedn_ctx;n_keepfor what to retain when tokens are discarded): https://raw.githubusercontent.com/ggml-org/llama.cpp/master/tools/server/README.md - Anthropic, Commercial Terms of Service ("Customer (a) retains all rights to its Inputs, and (b) owns its Outputs"; "may not train models on Customer Content") and its privacy centre on retention, read 24 September 2026, as one example of the words a provider's terms use, not as a statement of any provider's policy: https://www.anthropic.com/legal/commercial-terms and https://privacy.claude.com/en/articles/7996866-how-long-do-you-store-my-data
- Ollama library page for
qwen3.8:27b(default build: 18 GB, Q4_K_M model at 17 GB, 461M-parameter BF16 projector at 931 MB, 27.3B parameters): https://ollama.com/library/qwen3.8:27b - Qwen3-235B-A22B config (128 experts, 8 per token): https://huggingface.co/Qwen/Qwen3-235B-A22B
- Gemma 4 26B-A4B config (128 experts, top-8): https://huggingface.co/google/gemma-4-26B-A4B-it
- Wan 2.2 T2V-A14B card metadata: https://huggingface.co/Wan-AI/Wan2.2-T2V-A14B
- Licence and gating fields read via the Hub API on 20 September 2026 for: https://huggingface.co/meta-llama/Llama-4-Scout-17B-16E-Instruct , https://huggingface.co/black-forest-labs/FLUX.1-dev , https://huggingface.co/black-forest-labs/FLUX.2-dev , https://huggingface.co/black-forest-labs/FLUX.1-schnell , https://huggingface.co/black-forest-labs/FLUX.2-klein-4B , https://huggingface.co/stabilityai/stable-diffusion-3.5-large , https://huggingface.co/mlx-community/Qwen3.8-27B-4bit , https://huggingface.co/mlx-community/Qwen3.8-27B-8bit
- Hugging Face, safetensors repository and format specification (8-byte header length, JSON header, no code execution; pickle background): https://github.com/huggingface/safetensors
- DataCamp, safetensors joining the PyTorch Foundation in April 2026 (secondary; the move is reported here, not verified against a foundation announcement): https://www.datacamp.com/blog/safetensors-format
- GGUF specification (single-file format, metadata plus tensors, successor to GGML/GGMF/GGJT): https://github.com/ggerganov/ggml/blob/9c2adc4962a3a5d259f10db2171e0df5c83e4b05/docs/gguf.md and the Hugging Face Transformers GGUF page: https://huggingface.co/docs/transformers/en/quantization/gguf.md
- GGUF introduction date, August 2023 (Wikipedia, secondary; the llama.cpp pull request that merged the spec is dated November 2023): https://en.wikipedia.org/wiki/GGUF and https://github.com/ggml-org/ggml/pull/302
- Hugging Face Hub documentation, gated models (automatic and manual approval): https://huggingface.co/docs/hub/models-gated
- Hu et al., "LoRA: Low-Rank Adaptation of Large Language Models", arXiv 2106.09685 v2 (frozen weights, injected low-rank matrices, up to 10,000 times fewer trainable parameters on GPT-3): https://arxiv.org/abs/2106.09685v2
- Dettmers et al., "QLoRA: Efficient Finetuning of Quantized LLMs" (frozen 4-bit base, LoRA adapters, 65B on a 48 GB GPU): https://arxiv.org/abs/2305.14314v1
- Apache License 2.0 as shipped in the Qwen3.8-27B folder (clauses 4(b), 4(c), 4(d) and 6 for notices, modified files, NOTICE files and trademarks)
- Open Source Initiative, The Open Source AI Definition 1.0 (data information, code, parameters; adopted 27 October 2024): https://opensource.org/ai/open-source-ai-definition and board minutes https://opensource.org/meeting-minutes/2024-10-27
- Nelson F. Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni and Percy Liang, "Lost in the Middle: How Language Models Use Long Contexts" (Transactions of the Association for Computational Linguistics, volume 12, 2024, pages 157 to 173; arXiv preprint 2023; performance is often highest when relevant information is at the beginning or end of the input and significantly degrades when it is in the middle of a long context): https://arxiv.org/abs/2307.03172
NexusTier 03, language models (all read via the Hub API on 20 September 2026: licence and gating fields, safetensors parameter counts, config.json context and expert settings; card text where quoted)
- https://huggingface.co/Qwen/Qwen3.8-27B (thinking mode, reasoning_effort, named runtimes)
- https://huggingface.co/Qwen/Qwen3.8-Flash-Next
- bartowski's GGUF quantisations of Qwen3.8-27B, card read 24 September 2026 (perplexity and KL divergence of each quantisation against the BF16 reference): https://huggingface.co/bartowski/Qwen3.8-27B-GGUF
- llama.cpp, perplexity tool README, read 24 September 2026 (what perplexity measures; KL divergence of a quantised model's logits against an FP16 reference, 0 meaning identical): https://github.com/ggml-org/llama.cpp/blob/master/tools/perplexity/README.md
- https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash and https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro-0813
- vLLM documentation, "Parallelism and Scaling", read 24 September 2026 (tensor parallelism "if the model is too large for a single GPU"; the
GPU KV cache sizeandMaximum concurrencylines at startup): https://docs.vllm.ai/en/latest/serving/parallelism_scaling.html - Databricks, "LLM Inference Performance Engineering: Best Practices", read 24 September 2026 (decode speed set by loading parameters from memory; batching "increases throughput compared to processing queries sequentially, but each query will take longer to complete"; the cache grows with the number of sequences and their lengths): https://www.databricks.com/blog/llm-inference-performance-engineering-best-practices
- Rouhani et al., "Microscaling Data Formats for Deep Learning", read 24 September 2026 (MX formats "combine a per-block scaling factor with narrow floating-point and integer types"): https://arxiv.org/abs/2310.10537
- https://huggingface.co/moonshotai/Kimi-K3
- https://huggingface.co/zai-org/GLM-5.3
- https://huggingface.co/meta-llama/Llama-4-Scout-17B-16E-Instruct (17B activated, 10M context, early fusion, system protections) and https://huggingface.co/meta-llama/Llama-4-Maverick-17B-128E-Instruct (17B activated, 1M context)
- https://huggingface.co/google/gemma-4-26B-A4B-it (25.2B total, 3.8B active, 8 active of 128 plus 1 shared, 256K context per card) , https://huggingface.co/google/gemma-4-12B-it , https://huggingface.co/google/gemma-4-E4B-it , https://huggingface.co/google/gemma-4-E2B-it (effective-parameter naming, thinking token budgets, language counts, audio on E2B/E4B/12B)
- https://huggingface.co/openai/gpt-oss-120b and https://huggingface.co/openai/gpt-oss-20b (reasoning effort, chain of thought not for end users, MXFP4 and the 80 GB / 16 GB floors)
- https://huggingface.co/mistralai/Mistral-Small-4-119B-2603 (119B, 6.5B activated per token, branded A6B; 128 experts 4 active; 256k context per card; config exposes 1,048,576 positions)
- https://huggingface.co/microsoft/Phi-4
- https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16
- Ollama library entry for
qwen3.8:27b(parameters object, draft count): https://ollama.com/library/qwen3.8:27b - LM Studio named as a GGUF tool on the secondary Wikipedia GGUF page cited under Tier 02
NexusTier 03, image generation (Hub API fields and card text, 20 September 2026)
- https://huggingface.co/black-forest-labs/FLUX.2-dev (32B rectified flow transformer)
- https://huggingface.co/black-forest-labs/FLUX.2-klein-4B (~13 GB VRAM, RTX 3090/4070 and above)
- https://huggingface.co/stabilityai/stable-diffusion-3.5-large (MMDiT; three text encoders: OpenCLIP-ViT/G, CLIP-ViT/L, T5-xxl) and https://huggingface.co/stabilityai/stable-diffusion-3.5-medium
- https://huggingface.co/stabilityai/stable-diffusion-xl-base-1.0
- https://huggingface.co/Qwen/Qwen-Image-2512 (text rendering claim) and https://huggingface.co/Qwen/Qwen-Image
- https://huggingface.co/Tongyi-MAI/Z-Image-Turbo (8 NFEs, sub-second claim)
- https://huggingface.co/tencent/HunyuanImage-3.0 (64 experts, 80B total, 13B active)
- https://huggingface.co/Efficient-Large-Model/Sana_1600M_1024px_diffusers
NexusTier 03, video generation (Hub API fields and card text, 20 September 2026)
- https://huggingface.co/Lightricks/LTX-2.5 (audio modes, revenue-tiered licence) and https://huggingface.co/Lightricks/LTX-2
- https://huggingface.co/Wan-AI/Wan2.2-T2V-A14B (80 GB statement) and https://huggingface.co/Wan-AI/Wan2.2-TI2V-5B (24 GB / RTX 4090 statement; 720p 24 fps; MoE in a video diffusion model)
- https://huggingface.co/tencent/HunyuanVideo-1.5 (14 GB minimum with offloading)
- https://huggingface.co/genmo/mochi-1-preview (22 / 42 / 60 GB statements)
- https://huggingface.co/zai-org/CogVideoX-5b
- https://huggingface.co/stepfun-ai/stepvideo-t2v
NexusTier 03, speech and audio (Hub API fields and card text, 20 September 2026)
- https://huggingface.co/openai/whisper-large-v3 (>5M hours, 30-second receptive field, 99 language codes counted in the card's front matter) and https://huggingface.co/openai/whisper-large-v3-turbo
- https://huggingface.co/distil-whisper/distil-large-v3
- https://huggingface.co/nvidia/parakeet-tdt-0.6b-v3 and https://huggingface.co/nvidia/canary-1b-v2
- https://huggingface.co/Qwen/Qwen3-ASR-1.7B
- https://huggingface.co/UsefulSensors/moonshine-base
- https://huggingface.co/pyannote/speaker-diarization-community-1 and https://huggingface.co/pyannote/segmentation-3.0
- https://huggingface.co/nvidia/diar_sortformer_4spk-v1 and https://huggingface.co/nvidia/Nemotron-3-Diarization-preview
- https://huggingface.co/hexgrad/Kokoro-82M
- https://huggingface.co/Qwen/Qwen3-TTS-12Hz-1.7B-Base
- https://huggingface.co/nari-labs/Dia-1.6B
- https://huggingface.co/sesame/csm-1b
- https://huggingface.co/kyutai/pocket-tts
- https://huggingface.co/coqui/XTTS-v2 , https://huggingface.co/SWivid/F5-TTS , https://huggingface.co/fishaudio/fish-speech-1.5 , https://huggingface.co/bosonai/higgs-audio-v2-generation-3B-base
- https://huggingface.co/stabilityai/stable-audio-open-1.0 and https://huggingface.co/google/magenta-realtime-2
NexusTier 03, vision understanding (Hub API fields and card text, 20 September 2026)
- https://huggingface.co/Qwen/Qwen3-VL-8B-Instruct and https://huggingface.co/Qwen/Qwen3-VL-32B-Instruct
- https://huggingface.co/facebook/sam3 and https://huggingface.co/facebook/sam2.1-hiera-large
- https://huggingface.co/openai/clip-vit-large-patch14 (contrastive training description) and https://huggingface.co/google/siglip2-so400m-patch16-512
- https://huggingface.co/google/vit-base-patch16-224
- https://huggingface.co/microsoft/Florence-2-large , https://huggingface.co/deepseek-ai/DeepSeek-OCR , https://huggingface.co/PaddlePaddle/PaddleOCR-VL
NexusTier 03, embeddings and rerankers (Hub API fields and card text, 20 September 2026)
- https://huggingface.co/Qwen/Qwen3-Embedding-0.6B (dimensions 32 to 1,024; 100+ languages; MTEB claim), https://huggingface.co/Qwen/Qwen3-Embedding-4B , https://huggingface.co/Qwen/Qwen3-Embedding-8B , https://huggingface.co/Qwen/Qwen3-Reranker-0.6B , https://huggingface.co/Qwen/Qwen3-VL-Embedding-2B
- https://huggingface.co/google/embeddinggemma-300m (768 dims, MRL 512/256/128, on-device)
- https://huggingface.co/jinaai/jina-embeddings-v5-text-small and https://huggingface.co/jinaai/jina-reranker-v3
- https://huggingface.co/BAAI/bge-m3 , https://huggingface.co/nomic-ai/nomic-embed-text-v2-moe , https://huggingface.co/vidore/colpali-v1.3 , https://huggingface.co/sentence-transformers/all-MiniLM-L6-v2 , https://huggingface.co/mixedbread-ai/mxbai-embed-large-v1
- Muennighoff et al., "MTEB: Massive Text Embedding Benchmark" (8 tasks, 58 datasets, 112 languages; leaderboard address): https://arxiv.org/pdf/2210.07316v2 and the Hugging Face MTEB blog post https://huggingface.co/blog/mteb
NexusTier 03, classical machine learning
- scikit-learn user guide, table of contents (supervised, unsupervised and semi-supervised sections; ensembles): https://scikit-learn.org/stable/user_guide
ApexTier 04, the decision-model box (all secondary or vendor; nothing verified against weights or a paper)
- DataCamp, "Jev: TypeSafe's System One Model" (announcement date, category name, input and output shape, Jevons naming): https://www.datacamp.com/blog/system-one-models-jev
- Runtime Wire, "TypeSafe opens Jev early access" (no weights or reproducible paper; RLCD named; Kahneman reference; speed claims vendor-tested): https://runtimewire.com/article/typesafe-jev-system-one-ai-model-early-access
- Pasquale Pillitteri, "TypeSafe launches Jev" (Sean Goedecke's analysis: prefill plus single-token decoding, reproduction attempts, can still choose wrongly, capability ceiling): https://pasqualepillitteri.it/en/news/16460/typesafe-jev-chatless-ai-beats-claude
- TypeSafe AI documentation, "Introduction" (the three question shapes and what each returns; questions over one state evaluated in parallel and in isolation; ask one well-scoped thing per question and combine in code): https://docs.typesafe.ai/introduction
- TypeSafe AI documentation, "Confidence" (confidence as the concentration of the returned distribution; Noul answers carry none): https://docs.typesafe.ai/confidence
- TypeSafe AI documentation, "System One" (trained for calibrated decisions; probabilities optimised against outcomes): https://docs.typesafe.ai/concepts/system-one
ApexTier 04, training
- TypeSafe AI documentation, "Machine learning primer" (RLCD named; calibration as outcomes at 0.8 occurring about 80 per cent of the time; the contrast with RLHF): https://docs.typesafe.ai/introduction/machine-learning-primer
- "Scaling Laws for Neural Language Models" (arXiv 2001.08361; the loss scales as a power law with model size, dataset size and training compute): https://arxiv.org/abs/2001.08361
- DeepSeek-R1 card (R1-Zero trained by RL without SFT; cold-start data; six distilled models; 32B claim): https://huggingface.co/deepseek-ai/DeepSeek-R1 and https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Qwen-32B
- Ouyang et al., "Training language models to follow instructions with human feedback" (InstructGPT, RLHF): https://arxiv.org/abs/2203.02155
- Bai et al., "Constitutional AI" (RLAIF): https://arxiv.org/abs/2212.08073
- Rafailov et al., "Direct Preference Optimization": https://arxiv.org/abs/2305.18290
- Shao et al., "DeepSeekMath" (GRPO): https://arxiv.org/abs/2402.03300
- DeepSeek-AI, "DeepSeek-R1" paper: https://arxiv.org/abs/2501.12948
- Phi-4 card (synthetic "textbook-like" data): https://huggingface.co/microsoft/phi-4
- Minitron-8B-Base card (pruning of Nemotron-4 15B, then distillation): https://huggingface.co/nvidia/Minitron-8B-Base
- Gemma 4 QAT repositories, e.g. https://huggingface.co/google/gemma-4-12B-it-qat-q4_0-gguf and https://huggingface.co/google/gemma-4-26B-A4B-it-qat-q4_0-unquantized
- scikit-learn user guide (supervised, unsupervised, semi-supervised sections): https://scikit-learn.org/stable/user_guide
ApexTier 04, architecture words
- Vaswani et al., "Attention Is All You Need" (arXiv 1706.03762, June 2017; the Transformer, an architecture based on attention alone): https://arxiv.org/abs/1706.03762
- Qwen3.8-27B card and config (hidden layout, 48 linear-attention and 16 full-attention layers, 24 query and 4 key-value heads, MTP): https://huggingface.co/Qwen/Qwen3.8-27B
- Kimi-K3 card and config (2.8T total, 104B activated, 896 experts, 16 selected, 2 shared, 1,048,576 context; the card names its MoE "Stable LatentMoE"): https://huggingface.co/moonshotai/Kimi-K3
- Latent MoE paper (NVIDIA, 2026): https://arxiv.org/html/2601.18089v1 ; Sebastian Raschka's architecture gallery entry (secondary): https://sebastianraschka.com/llm-architecture-gallery/latent-moe/
- Cosmos 3 Super card (Mixture-of-Transformers, autoregressive and diffusion towers): https://huggingface.co/nvidia/Cosmos3-Super
ApexTier 04, variants
- Abliterated Qwen3.8-27B repositories found by hub search on 20 September 2026, e.g. https://huggingface.co/huihui-ai/Huihui-Qwen3.8-27B-abliterated (Apache 2.0)
- Abliterated Llama 3.1 example carrying the
llama3.1licence field: https://huggingface.co/mlabonne/Meta-Llama-3.1-8B-Instruct-abliterated - Guard models: https://huggingface.co/meta-llama/Llama-Guard-4-12B and https://huggingface.co/meta-llama/Llama-Prompt-Guard-2-86M ; the Llama 4 Scout card's system-protections statement
ApexTier 04, bolt-ons and agents
- Anthropic, "Donating the Model Context Protocol and establishing the Agentic AI Foundation" (9 December 2025; co-founders; 10,000+ servers): https://www.anthropic.com/news/donating-the-model-context-protocol-and-establishing-of-the-agentic-ai-foundation
- OWASP, LLM01:2025 Prompt Injection (direct and indirect injection; a modified document altering an application's output), read 24 September 2026: https://genai.owasp.org/llmrisk/llm01-prompt-injection/
- Linux Foundation, launch of the Agent2Agent protocol project (23 June 2025; protocol created by Google): https://s24.q4cdn.com/538403808/files/doc_news/Linux-Foundation-Launches-the-Agent2Agent-Protocol-Project-to-Enable-Secure-Intelligent-Communication-Between-AI-Agents-2025.pdf
- AAIF founding projects including AGENTS.md (secondary): https://agentic-ai.readthedocs.io/en/latest/Standards/agentic-ai-foundation/
- Engineering-ladder attestations (all secondary; usage evidence only): https://medium.com/@neuraldev/loop-engineering-vs-graph-engineering-the-architecture-shift-quietly-reshaping-ai-agents-c83488435d23 , https://www.analyticsvidhya.com/blog/2026/08/agent-harness-loop-graph-engineering/ , https://towardsdatascience.com/graph-engineering-for-ai-agents-from-prompts-and-loops-to-workflows/ , https://github.com/ai-boost/awesome-harness-engineering , https://www.sitepoint.com/vibe-coding-2026-complete-guide/
ApexTier 04, words on the invoice
- OpenAI, prompt caching guide, read 20 September 2026 (whole rendered prefix must match; minimum cacheable length, read and write ratios and breakpoint controls are per model generation;
cached_tokensreporting): https://developers.openai.com/api/docs/guides/prompt-caching.md - OpenAI, batch endpoint reference (24-hour completion window is the only one supported; 50,000 requests per file): https://developers.openai.com/api/reference/cli/resources/batches/methods/create/index.md
- OpenAI, structured outputs guide (function calling or
json_schemaresponse format): https://developers.openai.com/api/docs/guides/structured-outputs - OpenRouter, structured outputs documentation (gateway example; per-model parameter support): https://openrouter.ai/docs/guides/features/structured-outputs
ApexTier 04, world models, robotics, 3D (Hub API fields and card text, 20 September 2026)
- https://huggingface.co/nvidia/Cosmos3-Super , https://huggingface.co/nvidia/Cosmos3-Nano , https://huggingface.co/nvidia/Cosmos3-Edge , https://huggingface.co/nvidia/Cosmos3-Nano-Policy-DROID
- https://huggingface.co/facebook/vjepa2-vitl-fpc64-256 and https://huggingface.co/facebook/vjepa2-vitg-fpc64-384
- https://huggingface.co/lerobot/pi0 and https://huggingface.co/lerobot/pi05_base
- https://huggingface.co/microsoft/TRELLIS.2-4B , https://huggingface.co/tencent/Hunyuan3D-2.1 , https://huggingface.co/facebook/sam-3d-objects , https://huggingface.co/facebook/sam-3d-body-dinov3
ApexTier 04, evaluation and provenance
- MTEB paper, "no particular text embedding method dominates across all tasks": https://arxiv.org/pdf/2210.07316v2
- EleutherAI, lm-evaluation-harness ("a unified framework to test generative language models"; "the backend for Hugging Face's popular Open LLM Leaderboard"), read 24 September 2026: https://github.com/EleutherAI/lm-evaluation-harness
- Evaluation glossary showing the regression/drift/decay disagreement (secondary): https://futureagi.com/blog/llm-evaluation-glossary-definitions-2026/
- Google Cloud documentation, Content Credentials (C2PA description; validation failure on non-C2PA modification): https://docs.cloud.google.com/vertex-ai/generative-ai/docs/content-credentials
- European Commission, Code of Practice on Transparency of AI-Generated Content: press release of 10 June 2026 (obligations apply from 2 August 2026) https://digital-strategy.ec.europa.eu/en/news/commission-publishes-code-practice-marking-and-labelling-ai-generated-content and FAQ https://digital-strategy.ec.europa.eu/en/faqs/signing-code-practice-transparency-ai-generated-content
- Google DeepMind, SynthID technology page (watermarking and identification across images, audio, text and video; beta): https://deepmind.google/technologies/synthid/
- Further SynthID and provenance reading (secondary): https://compliancehub.wiki/eu-ai-act-marking-labelling-code-of-practice-article-50-2026/ , https://c2paviewer.com/articles/eu-ai-content-labels-c2pa , https://c2paviewer.com/articles/verify-ai-generated-image-c2pa-synthid , https://mvidmar.substack.com/p/ai-watermark-illusion-en
ApexTier 04, reading the table
- "Measuring Massive Multitask Language Understanding" (MMLU; arXiv 2009.03300; 57 tasks from elementary mathematics to law): https://arxiv.org/abs/2009.03300
- "MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark" (arXiv 2406.01574; trivial questions removed, ten options): https://arxiv.org/abs/2406.01574
- "GPQA: A Graduate-Level Google-Proof Q&A Benchmark" (arXiv 2311.12022; 448 questions written by domain experts in biology, physics and chemistry): https://arxiv.org/abs/2311.12022
- "SWE-bench: Can Language Models Resolve Real-World GitHub Issues?" (arXiv 2310.06770; edit a codebase to resolve a described issue): https://arxiv.org/abs/2310.06770
- "Humanity's Last Exam" (arXiv 2501.14249; 2,500 questions across dozens of subjects, written by subject-matter experts): https://arxiv.org/abs/2501.14249
- "Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference" (arXiv 2403.04132; crowdsourced pairwise comparison): https://arxiv.org/abs/2403.04132
Corrections
None yet. Corrections are logged here with the date and the source that prompted them.