The Origin
For the person who has heard the words and has no picture.
The AI you met today
The intro took the phrase apart. This section is about the AI you use without choosing to, which is most of it, and most of it does not write. Before you opened a chat product this morning you had probably used several of these already. The filter that sorts spam out of your inbox sorts. The system that flags an odd card transaction scores. The arrival time on the map predicts. The grouping of your photos by face, the search that finds "beach" among them and the unlock that recognises you all recognise. The captions on a call listen. The thing that suggests the next video ranks. None of these makes anything new. Each takes something in (a message, a payment, a photo, a sound) and gives back a label, a number, or a yes or a no. All of these are gathered at 1.4 as the kind that sorts, scores, predicts or finds, and the Nexus tier gives them their proper names at 3.8. This is the everyday AI. It is what the phrase meant for most of its life, it is mostly small, it mostly runs on an ordinary processor or on the phone's own chip, and between them these things still do most of the AI work that is actually in use.
The writing kind is newer, and it is what the news is about. The keyboard that offers your next word is doing, in miniature, what a chat model does: guessing the next piece from the pieces so far. Translation now writes a new sentence rather than looking one up. The paragraph a search engine puts above its results is written by a language model from the pages it found. The tool that removes a passer-by from a photo invents what was behind them. The chat products, the helper in the document editor, the notes written up after a call and the assistant you speak to are all the writing kind, sometimes with a listening model in front and a voice model behind. The trade's word for all of this is generative: it makes something new. Nothing in the word says the new thing is true.
Whether a given feature is the sorting kind or the writing kind, and whether it runs on the device in your hand or on a computer somewhere else, is a property of the product, and the product's own pages say which; the next section gives you the questions to ask. This guide is about the writing kind, and about one part of it in particular, the model, because that part can be downloaded and looked at, and once you have seen one you can read every claim about the rest.
The AI you use without choosing to, which is most of it, and most of it does not write
makes nothing new
each takes something ina messagea paymenta photoa sound
- The filter that sorts spam out of your inboxsorts
- The system that flags an odd card transactionscores
- The arrival time on the mappredicts
- The grouping of your photos by facerecognises
- the search that finds "beach" among themrecognises
- the unlock that recognises yourecognises
- The captions on a calllisten
- The thing that suggests the next videoranks
and gives backa labela numbera yes or a no
None of these makes anything new. 1.4 gathers them as the kind that sorts, scores, predicts or finds, and these things still do most of the AI work that is actually in use.
makes something new
the trade's wordgenerative
- The keyboard that offers your next wordguessing the next piece from the pieces so far
- Translationwrites a new sentence rather than looking one up
- The paragraph a search engine puts above its resultswritten by a language model from the pages it found
- The tool that removes a passer-by from a photoinvents what was behind them
- The chat products, the helper in the document editor, the notes written up after a call and the assistant you speak tosometimes with a listening model in front and a voice model behind
It makes something new. Nothing in the word says the new thing is true.
Where do you meet one?
Almost everyone has used one of the writing kind already, behind a text box. What sits behind the box is three things: a model, a program reading it, and a website in front. Almost nobody meets the first of the three as what it is, a folder of files; the next section opens that folder. There are about six ways you meet one. What differs is what the front does, who operates the program, and whether the file is yours, someone else's, or not named at all.
A chat product. ChatGPT, Claude and Gemini are the names today, and the everyday word for any of them is a chatbot: a website or a phone app with a text box. Someone else holds the file, runs the program and built the front; you operate none of the three. What you type goes to their computer. What happens to it there, how long it is kept, who may read it, whether it is used as an example when a later version is trained, is governed by their terms and not by anything in the model; the Vector tier asks that question for all six ways at 2.2. This is where most people start. A voice mode is the same front on the sound row of the map; a research mode is a search the program runs before answering; a copilot, as a common noun, is the second way below. Everything the front shows you is a property of the program or the front and not of the file: a picker offering a choice of models, one of them perhaps marked as thinking, a cap on how much you may use, a memory of earlier chats. The picker is at 2.11, thinking at 2.4, the cap at 1.6, the memory at 1.8.
AI inside a program you already had. A button in a mail program that drafts the reply, a helper in a document editor, the grey text that finishes your line in a code editor, notes written up after a call. The same three parts, but the front is a feature of a program made for something else, and you may not have chosen to use a model at all. This is not the spam filter of 1.1: that one sorts, this one writes.
A coding agent. A program that reads a folder of code, edits the files, runs commands, and usually asks before the risky ones; Claude Code, Codex and Cursor are names you will hear. The file may be someone else's or one you run yourself; the program is on your machine, or on a server that hands back the finished change. The plain word is the estate agent's: someone you give a goal to, who takes the steps on your behalf and decides some of them without asking you. The Apex tier calls a model in a loop with tools and a goal an agent, calls the program around it a harness, and has the rest of that vocabulary. Here the shape is enough: the program runs the loop and holds the permissions and the tools (the other programs it may call), and the model chooses the next step inside it.
An agent that runs without you. The same shape pointed at something other than code: a program that drives a browser or a screen, fills in forms, or answers customers, on its own or on a schedule. It is the three parts with the loop given permission to act, and the Apex tier says what the permissions are called and what goes wrong when the text it reads tells it to do something else.
An API. An API is a way for one program to ask another to do something over a network, with no screen and no person in between. Here a program of your own sends text to a model that a runtime somewhere is answering requests for, rather than a person typing into a front; the Vector tier calls that serving. The model may be one you operate or one a provider does; the API is only the way programs reach it. Where it is a provider's, use is typically metered, and for a language model the meter counts tokens. The front, if there is one, is yours to build. The Vector tier says what a token is and what "API-only" means; this guide does not say how to call one.
A download. The folder from 1.3 and a program on your machine to read it: Ollama, LM Studio and llama.cpp are the ones this guide names. You operate the file and the program, and the front, where there is one, runs beside them. This is the arrangement the guide is written from, and not because it is the common one: here the file and the program are in front of you rather than behind a product. Every name in a picker, every cap on a plan and every claim on a product page is a property of a file or of a program, and this is where you can see both plainly.
They are not six kinds of thing: a coding agent can sit on any of the others.
Who controls each part
- yoursyou control it
- theirssomeone else does
- eitherdepends on the settings
Three labels will follow you through the guide, and they answer two questions. Local, or on-device, says where: the program that reads the file, which this guide calls the runtime from here on, is on the device in front of you. Self-hosted and provider-hosted say who operates it: you, or somebody else; provider-hosted is what everyday speech calls the cloud, someone else's computer reached over the internet. A chat product is provider-hosted. A download is self-hosted, and local when it is on the machine you are sitting at. An API and an agent may be any of these. The Vector tier defines them properly.
The rest of this guide takes the file apart. Where a word belongs to something you have already done in a product, a box marked "You have met this" says so.
Why is it a download and not an app?
You will never need to open one of these yourself; this is what one looks like when someone does, and it is the picture the rest of the guide rests on. On the day this guide was checked, I opened the public folder for a model called Qwen3.8-27B. It holds 18 files named model-00001-of-00018.safetensors to model-00018-of-00018.safetensors, a config.json of about four kilobytes, a vocab.json, a chat_template.jinja, a LICENSE and a README.md. The 18 big files add up to 55.6 GB. A gigabyte, GB, is a thousand million bytes, and a byte is the unit files are measured in: eight bits, a bit being a single yes or no, the smallest thing a computer stores. So 55.6 GB is fifty-five thousand million bytes, and every one of them is part of a number. The index file that describes them lists 1,199 named blocks of numbers inside.
That is the whole release: a snapshot of the numbers as they stood when training stopped, which the trade calls a checkpoint. There is no .exe, no installer, nothing you double-click. The 55.6 GB is numbers. The four-kilobyte config is a description of how those numbers are arranged. A separate program, which the folder does not contain, reads both and does the work. The README names some of those programs: Transformers, vLLM, SGLang.
So "downloading a model" is a real sentence in a way "downloading ChatGPT" is not. ChatGPT is a product: a model, a program reading it, and a website in front. A model is only the first of those three. This is why the file can be enormous and the program that reads it can be small, and why the same file can be read by several different programs at different speeds.
Hold on to that picture. Almost every word in this guide is a property of that folder or of the program reading it.
Same word, different thing: model. In this guide, model usually means the released checkpoint (the folder of learned numbers, the weights, with the small files beside them) rather than the product wrapped around it. In a product announcement "the model" often means the file, the runtime and the service together. When someone says a model "got worse", ask which of the three they mean.
What the folder holds
model-00001-of-00018.safetensors onward55.6 GBthe numbersconfig.jsonabout 4 KBa description of how those numbers are arrangedvocab.json · chat_template.jinja · LICENSE · README.mdsmallWhat it does not hold
Why "downloading a model" is not "downloading ChatGPT"
- a website in front
- a program reading it
- the modelthis is the download
What does "AI" actually mean on the box?
The spam filter, the card check and the next-video suggestion from 1.1 are all sold as AI. None of them writes sentences or draws pictures.
Three boxes, one inside the other. Artificial intelligence is the outer box, the phrase the intro took apart: any program that does something we used to think needed a person, a boundary that moves as programs get reliable. Machine learning is inside it: programs that learn the behaviour from examples rather than being given rules. The examples are what the trade calls training data, and 1.8 says what learning from them does to the numbers. Deep learning is inside that: machine learning built on neural networks. A neural network is a long chain of simple arithmetic arranged in repeated stages called layers, with the numbers as the learned part; it is named after nerve cells because the earliest designs imitated one, and the likeness is loose, since nothing in the file resembles a brain except the idea of many simple parts joined up. Nearly every model in this guide is one. A deep network can be small, and the ones here happen to be large.
The split that matters more for you is a different one. Generative models make new things: text, pictures, sound, video. That is all the phrase generative AI means, and new does not mean true. The text-making kind is what the trade calls a large language model, LLM: a model that predicts the next piece of text from everything before it, which is the loop at 1.7, and large because its learned numbers run to billions, which the Vector tier counts at 2.5. It is the kind most of this guide is about. The letters GPT in ChatGPT expand to generative pre-trained transformer: generative as above, pre-trained as at 1.8, and transformer the name of the arrangement of the arithmetic, which the Apex tier explains at 4.2. Everything else sorts, scores, predicts or finds, which is the everyday AI of 1.1. The first kind is what the news is about and what this guide spends most of its time on. Keep both in view; the second is your anchor when the first gets strange.
Three boxes, one inside the other
- The spam filter
- the card check
- the next-video suggestion
What goes in and what comes out?
Ask a chat model a question and text comes back. Give a picture model a sentence and a picture comes back. Speak into a listening model and text comes back. Those three facts are the whole map.
Here it is in words, first drawing. Rows are what you give it, columns are what you get.
| In ↓ / Out → | Text | Picture | Sound | Video | A label or a number | A 3D shape or an action |
|---|---|---|---|---|---|---|
| Text | chatbot models | picture models | voice models | video models | sorting models | 3D and robot models |
| Picture | describing models | editing models | animating models | recognising models | 3D models | |
| Sound | listening models | voice-changing models | tagging models | |||
| Video | describing models | editing models | recognising models | |||
| Numbers | predicting models |
The names in the cells are deliberately plain. Every one of them has a proper name, and by the Nexus tier you will know them. For now the point is the shape of the table: a family is a pairing of an input and an output, and a good share of the confusion about AI comes from not knowing which pairing is being talked about.
You have met this: the text box. A chat product sits on the top-left cell and hides the rest of the table. Ask it for a picture and you have moved one cell to the right, which by this table is a different pairing and often a different model; talk to it and you have come in on the sound row. One front, several pairings, and the front does not show the seam.
This table is drawn once. From the next tier on, a table of the families grows beside it, tier by tier.
Rows are what you give it · columns are what you get
make new thingssorts, scores, predicts or findsthe whole map
- Text → TextAsk a chat model a question and text comes back.
- Text → PictureGive a picture model a sentence and a picture comes back.
- Sound → TextSpeak into a listening model and text comes back.
What is it actually reading?
What you type into the box has a name, the prompt. The plain word is for whatever moves someone to speak, as the prompter in the wings gives the actor the line. The model's prompt is more than what you typed: the product puts its own instructions and the conversation so far around your message before the model sees any of it, and the Vector tier shows that wrapping at 2.9.
I gave the sentence "Infrastructure architects rarely sleep." to the Qwen3.8-27B model's reader, its tokenizer, on the day this guide was checked. Four words, thirty-nine characters. The model received five pieces: Infrastructure, architects, rarely, sleep, .
The language model does not receive raw text, because a computer does arithmetic and nothing else. It receives a list of numbers, one per piece, and the pieces are called tokens, in the plain sense of the word: a counter that stands for something and is exchanged for it at the gate. Each piece of text is exchanged for a number that is only its place in the model's fixed list of pieces, the vocab.json in the folder at 1.3, and the number carries no meaning of its own. A common word is usually one token. A long or rare word is several. Punctuation is often its own token. With this tokenizer the word "strawberry" arrives as three pieces: str, aw, berry. For English text OpenAI's rule of thumb is that one token is about three-quarters of a word, or about four characters.
Other languages cost more, with this tokenizer; another model's tokenizer would cut the same sentences differently. The Arabic sentence for "the architect rarely sleeps", five words, arrived as ten tokens. The Urdu sentence for the same idea, eight words, arrived as twenty-three. A sixteen-digit number arrived as sixteen tokens, one per digit.
This matters because in the language world many of the limits, speeds and usage-based prices you will meet are expressed in tokens; the other families count differently, and the Nexus tier says how. The most a model can be given at once, everything you typed, everything it has written so far and whatever the product adds, is its context window, and it is a token count. How fast it writes is tokens per second. What a provider-hosted model bills you is tokens in and tokens out. None of those numbers is in words.
You have met this: the cap. With a plan that lets you use a model for a monthly fee, you are paying for it to be run on your behalf, and the product may meter you however it chooses: messages, credits, tasks, pictures, minutes of audio. Underneath, a language model's work is counted in tokens, which is the count above. The message that says you have reached your limit is a count of something, and the product's own pages say what is being counted.
It is also one reason for a famous failure. Ask a model how many r's are in "strawberry" and it may get it wrong; it received str, aw and berry, not letters, and has to reconstruct the spelling from what it learned.
What you type
Infrastructure architects rarely sleep.
What the model receives
And the reason it miscounts letters
What the same idea costs elsewhere
- The sentence above5 tokens1.25 / word
- Arabic · "the architect rarely sleeps"10 tokens2.00 / word
- Urdu · the same idea23 tokens2.88 / word
- A sixteen-digit number16 tokensone per digit
What does it do when it runs?
Here is what the model does with those pieces. From all the pieces so far it gives every piece in its list a score for being the next one; the program picks one, favouring the high scorers, adds it to the end, and asks again. That loop, run until the model produces its stop piece, is the whole of a chat, and it is why the answer appears a word or so at a time.
That loop has a name you will meet on invoices and in meetings: inference. In plain English an inference is a conclusion reached from evidence, what a detective does when she sees mud on your boots and concludes that you crossed the field; nobody told her, she worked it out from what was in front of her. The trade uses the word for everything a finished model does. From the pieces so far, and from what training left in its numbers, the model infers the next piece, and nothing in the file changes while it does so. Training made the file; inference uses it; every message you have ever sent to a chat product was inference. The likeness to the detective stops at one point. She can tell you her chain of reasoning. The model's is arithmetic over billions of numbers, and even when a model writes its working before its answer, which the Vector tier explains at 2.4, the working is more text from the same loop, not a report on the arithmetic.
Four things you have already seen follow from it. The answer streams onto the screen because it is being written one piece at a time, and how fast it writes is counted in pieces per second, which the last section called tokens per second. The same question can get two different answers because the pick is a draw weighted by those scores rather than always the single highest; the Nexus tier has the dials that shape the draw. A model in a thinking mode is slower because it writes working before its answer, and the working is pieces too, produced by the same loop; the Vector tier has a section on it. And a wrong answer can arrive as fluently as a right one, because the loop has no step at which it checks anything against a fact; the next section says why.
The loop, and its name: inference
- the modelgives every piece in its list a score for being the next onereading the file nothing in the file changes while it does so
- the programpicks one, favouring the high scorers
- the programadds it to the end
no step herechecks anything against a fact
Four things that follow from it
- the loopThe answer streams onto the screenbecause it is being written one piece at a timecounted in tokens per second
- step 2, the pickThe same question can get two different answersbecause the pick is a draw weighted by those scores rather than always the single highestthe dials that shape the draw
- the loopA model in a thinking mode is slowerbecause it writes working before its answerthe working is pieces too, produced by the same loop
- the check that is not therea wrong answer can arrive as fluently as a right onebecause the loop has no step at which it checks anything against a fact
How did it learn?
The Qwen3.8-27B card lists its training stage as "Pre-training & Post-training". That is the whole life story of a model in one line, and one paragraph is enough for now.
Training is the athlete's word, and it means here what it means on the track: practice with correction, repeated until the performance improves. A model starts as random numbers. It is shown an example, text for a chat model, a picture with its caption for a picture model, and asked to guess the next piece, the loop of the last section. It is told how wrong the guess was, and every one of its numbers is nudged a little in the direction that would have made it less wrong. Then the next example, and the next, in enormous batches, for weeks or months on many machines. Nobody writes a rule. Nobody tells it what a noun is. After enough examples the numbers encode the patterns. Post-training is a second, shorter phase where the model is shown examples of the behaviour people want from it, answering questions rather than continuing a paragraph, and nudged again.
Two consequences follow from this and both are words you will hear. First, the model has no fact database that ordinary generation consults; what it knows is encoded diffusely in its learned numbers, and that can produce a fluent, confident sentence that is wrong. The word for that is hallucination, borrowed from seeing what is not there; some in the trade prefer confabulation, the word for a sincere false account, since nothing was seen. Post-training, which nudges the model towards the answers people preferred, does not on its own reward it for saying "I do not know". Second, the model's examples stopped arriving on some date, and its numbers are not reliably informed by anything after it unless later post-training added it. The word for that is the knowledge cutoff. It can still be handed newer information at run time, pasted into what it reads or fetched by a search the program runs; the Apex tier calls that retrieval. Neither is a bug in one product. Both are what "learned from examples" means.
You have met this: it remembers me. When a product seems to know what you said last week, or knows about something after its cutoff, the file has not learned it; nothing in the file changes when it is used, which is what inference at 1.7 means. The program is handing it something at run time: what it kept from earlier chats, a search it ran, a document you gave it. The Apex tier calls the first memory and the second retrieval. Using it does not retrain it; the product is supplying the information again each time. A later version of the model may know it, and that is a different checkpoint, not this one having learned.
Same word, different thing: the borrowed words. Learn, know, understand, think, reason, remember and hallucinate are words for what a person does, and the trade uses every one of them for what a model does. Each is borrowed, and each means less than it sounds. It learned means its numbers were nudged on examples. It knows means the pattern is somewhere in those numbers. It remembers means the program handed it the earlier text again. It thinks means it wrote working before its answer. It hallucinated means the loop produced a fluent sentence that is not so. Whether any of these is the same as what you do is an argument this guide leaves to philosophy. When you hear one of these words about a model, swap in the mechanism and see whether the sentence still stands.
Pre-training & Post-training
- textfor a chat model
- pictures with captionsfor a picture model
- examples of the behaviour people wantanswering questions rather than continuing a paragraph
Nobody writes a rule.Nobody tells it what a noun is.
What follows from it
- hallucinationNo fact database that ordinary generation consults. What it knows is encoded diffusely in its learned numbers, which can produce a fluent, confident sentence that is wrong.
- the knowledge cutoffIts examples stopped arriving on some date, and its numbers are not reliably informed by anything after it, unless later post-training added it.It can still be handed newer information at run time, pasted into what it reads or fetched by a search the program runs; the Apex tier calls that retrieval.
What does it run on?
Every computer has a processor, the CPU, which runs ordinary programs, and main memory, the RAM, where a running program keeps what it is working on. A graphics card is a second processor, the GPU, with hundreds of small cores, thousands on current cards, that do the same simple arithmetic across a great many numbers at once. It carries its own memory, the VRAM, separate from the machine's RAM. The difference is in the shape of the work: the CPU has a few cores that will take on any job in turn, and the GPU has hundreds that do one kind of job at once. Those four words account for most of the hardware talk you will meet.
Graphics cards were built to draw pictures, where every point on the screen is worked out at the same time. A model's arithmetic has the same shape: the same sums over millions of numbers at once. That is why a chip made for games became the chip for AI. Most of the software that runs models was written for NVIDIA's cards first, through NVIDIA's own programming platform, CUDA; other makers' cards are reached by other routes, added later.
The difference is in the shape of the work
- any job
- then the next
- then the next
Two jobs with the same shape
- Graphics cards were built to draw picturesevery point on the screen is worked out at the same time
- A model is a file of numbersthe same sums over millions of numbers at once
That is why a chip made for games became the chip for AI.
A model is a file of numbers, and running it means doing that arithmetic again and again, for everything it reads and everything it writes. So the file has to be in memory while it runs. In the card's VRAM it runs at the card's speed. Where it does not fit, a runtime can keep part of the file in RAM and carry each part across at the moment it is needed, which is called offloading and costs time every time the model produces a piece of output. Where it does not fit in RAM either, a runtime can spill to disk, and the runtimes' own documentation calls that very slow.
Some machines keep one pool of memory that the processor and the graphics share, called unified memory, and Apple silicon Macs are the common case. There "fit" means fit in that pool, so a Mac with a lot of memory holds a file that would otherwise need a data-centre card, and the runtimes treat it as a first-class target. That is why people say Macs are good at this.
Small models run on a processor alone, more slowly, and the small chip now sold inside phones and laptops as an NPU is built for those; it is where the makers say much of the everyday kind from the start of this tier runs, on the phone itself: the keyboard's next word, the photo grouping, the unlock. The larger the file, the more the graphics card matters. Every memory figure later in this guide asks the same question. Does the file fit in the memory of the thing that will run it?
Two shapes of machine
Where the file sits, and what that costs
- In the card's VRAM it runs at the card's speed
- Part of the file in RAM offloadingcosts time every time the model produces a piece of output carry each part across at the moment it is needed
- Spill to disk the runtimes' own documentation calls that very slow
What you now know
If you stop here, this is what you can say, and each sentence is one you could defend.
Artificial intelligence is the name of a field and its hope, a label on products, and in this guide a model: a file of numbers learned from examples. Most of the AI in your day sorts, scores, predicts or finds, is small, and has been there for years. The writing kind is newer, is what the news is about, and is called generative because it makes something new, which is not the same as true.
A model is a large file of learned numbers with small files beside it that describe their shape. A product is that file, a program reading it, and a front. What you type goes wherever the program is, under the terms of whoever runs it.
Training is practice with correction: the model guesses the next piece of an example, is told how wrong it was, and its numbers are nudged, billions of times over, until they hold the patterns. The maker does it once, and then it is over. Inference is the detective's word, a conclusion from evidence, and it is everything the finished model does: from the pieces so far it infers the next, and nothing in the file changes.
A prompt is what you send, including what the product adds around what you typed. A token is a piece of text swapped for a number that is only its place in a fixed list, and everything the language world counts, it counts in tokens. The context window is the most tokens the model can be given at once. A parameter is one of the learned numbers, and "27B" counts them.
A hallucination is a fluent sentence that is not so, produced by the ordinary loop, because nothing in the loop checks against a fact. The knowledge cutoff is the date the examples stopped. When a product remembers you or knows today's news, the program handed it that at run time; the file did not learn it. An agent is a model in a loop with a goal and tools, inside a program that holds the permissions. A chatbot is the everyday word for a chat product. The cloud is someone else's computer.
Everything else in this guide is a property of the file or of the program reading it.
Every word, on the part it belongs to
artificial intelligencethe name of a field and its hopea label on productsand in this guide a model: a file of numbers learned from examples
A product is that file, a program reading it, and a front.
the file
- a modela large file of learned numbers with small files beside it that describe their shape
- trainingpractice with correctionThe maker does it once, and then it is over
- a parameterone of the learned numbers"27B" counts them
- a tokena piece of text swapped for a number that is only its place in a fixed list
- the context windowthe most tokens the model can be given at once
- the knowledge cutoffthe date the examples stopped
the program reading it
- inferenceeverything the finished model doesnothing in the file changes
- where what you type goeswherever the program is, under the terms of whoever runs it
- it remembers methe program handed it that at run timethe file did not learn it
- an agenta model in a loop with a goal and tools, inside a program that holds the permissions
- the cloudsomeone else's computer
the front
- a promptwhat you send, including what the product adds around what you typed
- a chatbotthe everyday word for a chat product
a hallucinationa fluent sentence that is not so: produced by the ordinary loop, because nothing in the loop checks against a fact. It sits on the loop, which is the program running the file, and so on neither part alone.
1.11The question Origin leaves you with
If a model is a 55.6 GB file of numbers, what are the numbers, why are there 27 billion of them, and why does the same model come in a dozen different sizes on the download page? That is the Vector tier.