From the last vector to the next word
Logits, softmax, temperature, top-k and top-p: how a model picks what to say next.
本文暂时只有英文版,网站的其他部分已翻译。
A language model reads your text as tokens, turns each one into a vector, and passes those vectors through a stack of transformer blocks (Attention and the transformer explains that part). At the end, every token has a refined vector. But the model is supposed to write something. This article follows the last step: from a list of numbers to an actual next word, and then the next, and the next.
One vector, one question
To predict what comes next, the model only needs the final vector of the last token. Thanks to attention, that vector has already gathered what it needs from everything before it. In GPT-2’s smallest version it is 768 numbers long.
The model’s vocabulary, the fixed list of tokens it knows (see Why models read tokens), has 50,257 entries in GPT-2. The job now is to give every one of those entries a score for “how well would this fit next?”
From vector to scores: the unembedding
The scoring is one multiplication. The model has a big learned table called the unembedding matrix (or output layer), with one column of 768 numbers per vocabulary token. Multiplying the last vector by this matrix is the same as taking a dot product with each column: multiply number by number and add up. A column that points in a similar direction to the vector gets a high score.
The result is one number per token, 50,257 in all. These raw scores are called logits. They can be any size and can be negative; only their differences matter. In GPT-2 the unembedding matrix is the same table as the embedding table from the very first step, reused (a trick called weight tying). Many newer models keep two separate tables.
From scores to probabilities: softmax
Logits are hard to use directly. What we want is a probability for each token: a number between 0 and 1, with all of them adding up to 1. The function that does this is softmax. For each logit it computes e raised to that logit (which makes every number positive and stretches the gaps), and then divides by the total.
For example, logits of 2.0, 1.0 and 0.1 become probabilities of about 65.9%, 24.2% and 9.9%. The biggest logit wins the most, but everything keeps a share. This list of probabilities is the model’s real output: not a word, but a probability distribution over every word it could say next.
Choosing a word: decoding
Turning that distribution into one token is called decoding, and there are several ways to do it. Try them below. The candidate words and their logits are made up for illustration, and to keep things readable we pretend the vocabulary has only these 12 words. The maths applied to them is the real thing.
Here is what each control does.
- Greedy decoding always takes the single most likely token. Set temperature to 0 to see it: “mat” gets 100% and every draw is the same. Greedy is predictable, but over long texts it tends to produce bland and repetitive writing.
- Temperature divides every logit by a number T before the softmax. Below 1 the gaps grow, so the favourite takes even more of the probability. Above 1 the gaps shrink and unlikely words get a real chance. It never changes the order of the words, only how sharply the probability is concentrated.
- Top-k keeps only the k most likely tokens and throws the rest away. It stops the model from ever picking something wildly unlikely from the long tail of a big vocabulary. Its weakness is that k is fixed: sometimes only two words make sense, sometimes hundreds do.
- Top-p, also called nucleus sampling, keeps the smallest group of top tokens whose probabilities add up to at least p. When the model is confident, that group is tiny; when many words fit, it grows. Set top-p to 0.50 above and only “mat” and “floor” survive, because “mat” alone is 41.3%, just short of half.
| Temperature | logit 2.0 | logit 1.0 | logit 0.1 |
|---|---|---|---|
| T = 0.5 | 86.4% | 11.7% | 1.9% |
| T = 1 | 65.9% | 24.2% | 9.9% |
| T = 2 | 50.2% | 30.4% | 19.4% |
Why the same question gets different answers
Unless decoding is greedy, the model makes a random draw at every single token. A 25% word will be picked about one time in four. Once a different word is picked, everything after it is predicted from a different text, so small early differences grow into completely different answers. That is why asking a chatbot the same question twice can give two different replies, and why settings like temperature exist: they let you trade variety for predictability.
The loop: append and repeat
One pass through the model produces one token. To write a sentence, the model adds the chosen token to the end of the input and runs again, predicting the token after that. This is called autoregressive generation: each step is fed its own previous output. It is also why replies appear on screen word by word.
Generation stops in one of two ways:
- An end-of-sequence token. The vocabulary contains a special token that means “this text is finished”. In GPT-2 it is
<|endoftext|>, ID 50256. The model learned during training that texts end with it, so it can predict it like any other token. When it is drawn, generation stops. - A length limit. The program running the model sets a maximum number of new tokens. If that is reached first, the text is cut off, sometimes mid-sentence.
Real systems add refinements, such as reusing earlier calculations so each step does not start from scratch, or searching several candidate continuations at once (beam search). But the core is exactly this: vector, logits, softmax, pick one, append, repeat.
Want to see the tokens a model actually works with? The Tokenizer on this site splits any text into GPT and LLaMA tokens with their IDs.
Further reading
- How to generate text: using different decoding methodsHugging Face blog · greedy, beam search, top-k and top-p with code
- Transformers, the tech behind LLMs3Blue1Brown · video, includes unembedding, softmax and temperature
- The Curious Case of Neural Text DegenerationHoltzman et al., 2019 · the paper that introduced nucleus (top-p) sampling
