← 学ぶ · ここから始めよう
クイックスタート8 分で読めます

How a language model turns your words into an answer

Text, tokens, vectors, a transformer, and one token at a time: the whole pipeline in plain language.

この記事は現在、英語のみです。サイトのほかの部分は翻訳されています。

You type “The cat sat on the” and a chatbot suggests “mat”. It feels like the model read your sentence, thought about it and answered. What actually happens is a fixed chain of simple steps that turns text into numbers, runs arithmetic on those numbers, and turns the result back into text. This article walks through that chain once, end to end, with no maths background assumed.

A language model is a program trained to continue text. A large language model (LLM) is the same idea at a huge scale: billions of adjustable numbers, trained on a large share of the text people have written. Everything below applies to the models behind ChatGPT, Claude, Gemini or Llama.

Entire LLM process, simplified

The promptembedding table~200,000 rows00000+ positionlast vector×unembedding matrix · one column per token=logits · one score for each of ~200,000 tokenseverything else · 27%top-k, k = 3

Start with a prompt.

0:00 / 1:19
One pass for “The cat sat on the” → “ mat”, then a second pass that adds “.”. Token IDs are real (o200k_base tokenizer, used by GPT-4o). The vectors, attention weights, layer count, scores and probabilities are illustrative, not taken from a real model.

The rest of the article takes those steps one at a time.

1. Text becomes tokens

A model can’t work with letters directly, so the first step is a tokenizer: a separate program that cuts text into pieces called tokens. A token is often a whole common word, sometimes a piece of a word, a punctuation mark or a space. Rare words are built from several smaller pieces.

“unbelievable” at the start of a text · 3 tokensunbelievable
“ unbelievable” after a space · 1 token·unbelievable
The same word cut two ways by the o200k_base tokenizer. At the very start of a text it is rare, so it is built from three pieces. After a space, as it usually appears, it is common enough to be one token. The dot marks the space.

Why pieces, rather than letters or whole words, is a story of its own: see why models read tokens, not letters or words, and how tokenizers are built for the algorithm that picks the pieces.

2. Tokens become ID numbers

Every tokenizer has a fixed vocabulary: a numbered list of all the pieces it knows. For o200k_base that list has about 200,000 entries. Each token is replaced by its number in the list, its token ID. Our sentence becomes 976 9059 10139 402 290.

These numbers are only positions in a list, like seat numbers. Token 9059 (·cat) is not “bigger” or “more” than token 402 (·on) in any useful sense. The model needs something richer to work with.

3. IDs become vectors

Inside the model sits a huge table with one row per vocabulary entry. Each row is a vector: a list of numbers, typically a few thousand long in a large model. Looking up a token’s row gives its embedding. Nobody writes these numbers by hand. They start random and are adjusted during training, until tokens used in similar ways end up with similar rows.

The embedding table as a lookup: each token ID selects one row. Colours stand for numbers (teal positive, red negative). Illustrative: only 8 of the thousands of columns are drawn, and the values are made up.

That idea, meaning as position in a space of numbers, is the heart of what embeddings are, and latent space explores what the directions in that space can capture.

4. The transformer: tokens look at each other

At this point each vector describes its token alone. The word “bank” gets the same starting vector in “river bank” and in “bank account”. The main body of the model, the transformer, fixes that. It is a stack of identical-looking layers, dozens of them in a large model.

Each layer does two things. First, attention: every token looks back at the tokens before it, decides which ones matter for it, and mixes in information from them. Then a small neural network updates each vector on its own. After many layers, each vector describes its token in context.

The same token “·bank” enters the layers with an identical vector in both sentences. After attention has mixed in the surrounding words, the two vectors differ. Illustrative values.

Attention is the idea that made modern language models work so well. It gets a full, step-by-step treatment in attention and the transformer, explained from scratch.

5. Only the last position predicts

After the final layer there is still one vector per token. To choose the next token, the model uses only the vector at the last position, the one for ·the in our example. Because attention let it gather information from every earlier token, that single vector now carries what the model has worked out about the whole sentence.

6. A score for every token: logits

The last vector is compared against every entry in the vocabulary, producing one score per token: about 200,000 numbers. These raw scores are called logits. A higher logit means the model rates that token as a better continuation. On their own, logits are hard to read: they can be negative and they don’t add up to anything in particular.

7. Scores become probabilities: softmax

A function called softmax turns the logits into probabilities. It makes every value positive and scales them so they add up to 100%. Bigger logits get a much bigger share, because softmax exponentiates each score: two points more of logit means about 7 times more probability.

next tokenlogitprobability
·mat9.038%
·floor8.217%
·bed7.58%
·couch7.26%
·rug6.84%
everything elselower27%
Logits and the probabilities softmax makes from them, for the five most likely tokens after “The cat sat on the”. Illustrative numbers, chosen so they are consistent with each other; the remaining ~200,000 tokens share the last 27%.

8. Picking one token

Now the model has a probability for every token, and something has to choose. The simplest rule, greedy decoding, always takes the top one. Most chat systems instead sample: they draw at random, weighted by the probabilities, so the same prompt can give different answers.

Settings shape that draw. Top-k sampling keeps only the k most likely tokens and rescales them to 100% before drawing, so wildly unlikely tokens can never appear. Temperature makes the probabilities sharper or flatter. From the last vector to the next word shows each of these with interactive examples.

9. Append and repeat

The chosen token, ·mat (ID 2450), is added to the end of the text. Then the whole pipeline runs again on the longer text to choose the token after it, and again after that. A reply of 300 tokens is 300 passes. This loop is why answers appear on screen word by word.

pass 1 reads 5 tokens → adds ·matThe·cat·sat·on·the
pass 2 reads 6 tokens → adds .The·cat·sat·on·the·mat
pass 3 reads 7 tokens → …The·cat·sat·on·the·mat.
Each pass reads everything so far and adds one token. The loop stops when the model produces a special end-of-text token or reaches a length limit.

Real systems avoid redoing all the work each pass by saving intermediate results for tokens they have already processed, but the logic is the same: one pass, one new token.

Where the numbers come from

Every number the model uses, the embedding table, the attention layers and the final scoring, is a parameter learned in training. Training shows the model enormous amounts of text, asks it to predict each next token, and nudges all the parameters slightly so the true next token becomes a little more likely. Repeated over trillions of tokens, those nudges add up to a model that continues text well. Chat assistants get further training on conversations so their continuations read as helpful answers.

Common misconceptions

  • “It writes the whole answer at once.” It doesn’t. It predicts a single token, adds it, and runs again. It has no finished answer waiting; each token is chosen given only the text so far.
  • “It looks the answer up.” There is no database of questions and answers inside. There is only the fixed set of learned parameters, and the answer is computed fresh from them every time. That is also why a model can state something false with confidence: nothing checks the result against a source.
  • “It understands like a person does.” What the model has are statistical patterns learned from text, rich enough to capture grammar, facts and some reasoning. Whether that counts as understanding is debated. What is certain is that it learned from text, not from seeing, touching or living in the world.

Where to go next

Each step above has its own article: tokens, tokenizers, embeddings, latent space, attention and choosing the next token. On this site you can split your own text with the Tokenizer and explore word vectors in the Embedding explorer.

Further reading