How a language model turns your words into an answer
Text, tokens, vectors, a transformer, and one token at a time: the whole pipeline in plain language.
この記事は現在、英語のみです。サイトのほかの部分は翻訳されています。
You type “The cat sat on the” and a chatbot suggests “mat”. It feels like the model read your sentence, thought about it and answered. What actually happens is a fixed chain of simple steps that turns text into numbers, runs arithmetic on those numbers, and turns the result back into text. This article walks through that chain once, end to end, with no maths background assumed.
A language model is a program trained to continue text. A large language model (LLM) is the same idea at a huge scale: billions of adjustable numbers, trained on a large share of the text people have written. Everything below applies to the models behind ChatGPT, Claude, Gemini or Llama.
Entire LLM process, simplified
Start with a prompt.
The rest of the article takes those steps one at a time.
1. Text becomes tokens
A model can’t work with letters directly, so the first step is a tokenizer: a separate program that cuts text into pieces called tokens. A token is often a whole common word, sometimes a piece of a word, a punctuation mark or a space. Rare words are built from several smaller pieces.
Why pieces, rather than letters or whole words, is a story of its own: see why models read tokens, not letters or words, and how tokenizers are built for the algorithm that picks the pieces.
2. Tokens become ID numbers
Every tokenizer has a fixed vocabulary: a numbered list of all the pieces it knows. For o200k_base that list has about 200,000 entries. Each token is replaced by its number in the list, its token ID. Our sentence becomes 976 9059 10139 402 290.
These numbers are only positions in a list, like seat numbers. Token 9059 (·cat) is not “bigger” or “more” than token 402 (·on) in any useful sense. The model needs something richer to work with.
3. IDs become vectors
Inside the model sits a huge table with one row per vocabulary entry. Each row is a vector: a list of numbers, typically a few thousand long in a large model. Looking up a token’s row gives its embedding. Nobody writes these numbers by hand. They start random and are adjusted during training, until tokens used in similar ways end up with similar rows.
row = token ID · columns = numbers in the vector
That idea, meaning as position in a space of numbers, is the heart of what embeddings are, and latent space explores what the directions in that space can capture.
4. The transformer: tokens look at each other
At this point each vector describes its token alone. The word “bank” gets the same starting vector in “river bank” and in “bank account”. The main body of the model, the transformer, fixes that. It is a stack of identical-looking layers, dozens of them in a large model.
Each layer does two things. First, attention: every token looks back at the tokens before it, decides which ones matter for it, and mixes in information from them. Then a small neural network updates each vector on its own. After many layers, each vector describes its token in context.
I sat by the river bank
I opened a bank account
Attention is the idea that made modern language models work so well. It gets a full, step-by-step treatment in attention and the transformer, explained from scratch.
5. Only the last position predicts
After the final layer there is still one vector per token. To choose the next token, the model uses only the vector at the last position, the one for ·the in our example. Because attention let it gather information from every earlier token, that single vector now carries what the model has worked out about the whole sentence.
6. A score for every token: logits
The last vector is compared against every entry in the vocabulary, producing one score per token: about 200,000 numbers. These raw scores are called logits. A higher logit means the model rates that token as a better continuation. On their own, logits are hard to read: they can be negative and they don’t add up to anything in particular.
7. Scores become probabilities: softmax
A function called softmax turns the logits into probabilities. It makes every value positive and scales them so they add up to 100%. Bigger logits get a much bigger share, because softmax exponentiates each score: two points more of logit means about 7 times more probability.
| next token | logit | probability |
|---|---|---|
| ·mat | 9.0 | 38% |
| ·floor | 8.2 | 17% |
| ·bed | 7.5 | 8% |
| ·couch | 7.2 | 6% |
| ·rug | 6.8 | 4% |
| everything else | lower | 27% |
8. Picking one token
Now the model has a probability for every token, and something has to choose. The simplest rule, greedy decoding, always takes the top one. Most chat systems instead sample: they draw at random, weighted by the probabilities, so the same prompt can give different answers.
Settings shape that draw. Top-k sampling keeps only the k most likely tokens and rescales them to 100% before drawing, so wildly unlikely tokens can never appear. Temperature makes the probabilities sharper or flatter. From the last vector to the next word shows each of these with interactive examples.
9. Append and repeat
The chosen token, ·mat (ID 2450), is added to the end of the text. Then the whole pipeline runs again on the longer text to choose the token after it, and again after that. A reply of 300 tokens is 300 passes. This loop is why answers appear on screen word by word.
Real systems avoid redoing all the work each pass by saving intermediate results for tokens they have already processed, but the logic is the same: one pass, one new token.
Where the numbers come from
Every number the model uses, the embedding table, the attention layers and the final scoring, is a parameter learned in training. Training shows the model enormous amounts of text, asks it to predict each next token, and nudges all the parameters slightly so the true next token becomes a little more likely. Repeated over trillions of tokens, those nudges add up to a model that continues text well. Chat assistants get further training on conversations so their continuations read as helpful answers.
Common misconceptions
- “It writes the whole answer at once.” It doesn’t. It predicts a single token, adds it, and runs again. It has no finished answer waiting; each token is chosen given only the text so far.
- “It looks the answer up.” There is no database of questions and answers inside. There is only the fixed set of learned parameters, and the answer is computed fresh from them every time. That is also why a model can state something false with confidence: nothing checks the result against a source.
- “It understands like a person does.” What the model has are statistical patterns learned from text, rich enough to capture grammar, facts and some reasoning. Whether that counts as understanding is debated. What is certain is that it learned from text, not from seeing, touching or living in the world.
Where to go next
Each step above has its own article: tokens, tokenizers, embeddings, latent space, attention and choosing the next token. On this site you can split your own text with the Tokenizer and explore word vectors in the Embedding explorer.
Further reading
- But what is a GPT? Visual intro to transformers3Blue1Brown, YouTube · the best visual walk-through of this whole pipeline (now titled “Transformers, the tech behind LLMs”)
- The Illustrated GPT-2Jay Alammar · the same pipeline drawn step by step, one level deeper
- LLM VisualizationBrendan Bycroft · a 3D walk through every number in a small real GPT
