Why models read tokens, not letters or words
The trade-off between letters, words and word pieces, and why every modern model settled on pieces, with bytes as a safety net.
Por ahora, este artículo solo está disponible en inglés. El resto del sitio está traducido.
A language model is a program that does arithmetic. It cannot read the letters on your screen directly. Before it can do anything with your text, the text has to be chopped into small units and each unit swapped for a number. Those units are called tokens, and the chopping is called tokenization.
The obvious question is: what should a unit be? A single letter? A whole word? Every modern model answers “something in between”, and this article explains why. It comes down to one trade-off between how many different units you allow and how many of them it takes to write a sentence.
Two numbers that pull against each other
Whatever unit you pick, two numbers describe the result.
- The vocabulary is the fixed list of every unit the model knows. Each entry gets an ID, and the model learns a separate set of numbers for every entry. A bigger vocabulary means more to learn and store.
- The sequence length is how many units it takes to write a particular text. Longer sequences mean more work every time the model reads or writes.
Small units give a small vocabulary but long sequences. Big units give short sequences but a huge vocabulary. Try it below: the same text, cut three ways.
23 characters · 6 words · 7 tokens
Look at a few of the examples. Characters always give the longest row. Words give the shortest row for English, but they fall apart on Japanese, which is written without spaces, so the whole sentence becomes one “word”. The subword row is close to the word row in length, yet it copes with every example.
Option one: characters
Using one token per character has real appeal. The vocabulary is small and never runs out: any text is just a string of characters, so the model can never meet something it cannot write down.
The problem is length. A character carries very little meaning on its own. The letter t tells the model almost nothing, so the model has to spend its effort gluing letters back into words before it can think about what they mean. And every text becomes several times longer than it needs to be.
Length matters for two concrete reasons. First, compute: in a transformer, every token looks at every other token, so that part of the work grows with the square of the length. Double the length and that part costs about four times as much. Second, the context window, the maximum number of tokens a model can take in at once, is a fixed budget. If each token holds less text, less of your document fits.
“It is a truth universally acknowledged, that a single man in possession of a good fortune, must be in want of a wife.”
Option two: whole words
Whole words fix the length problem: in English, the token count is close to the word count. Each unit also means something. But three problems appear.
- The vocabulary explodes. English alone has hundreds of thousands of word forms once you count plurals, tenses and compounds (
run,runs,running,rerun…). Add names, numbers, typos, hashtags, code and other languages, and no list is ever complete. - Unknown words. Anything missing from the list has to be replaced by a single catch-all token, usually written
[UNK]. Every unknown word looks identical to the model, so its meaning is simply lost. - Not every language uses spaces. Chinese, Japanese and Thai don’t put spaces between words, so “split at spaces” doesn’t even define what a word is.
Rare words cause a quieter problem too. A word that appears only a handful of times in the training text gets its own vocabulary entry, but the model sees too few examples to learn much about it.
The compromise: subword pieces
Subword tokenization takes the useful half of each option. Common words, and common words with a space in front, get a single token of their own, so everyday English stays short. Rare words are spelled out from a few smaller pieces that are themselves common: in the example above, Tokenizers becomes Token + izers. The pieces are reusable, so the model learns what izers tends to mean from many different words.
The vocabulary stays at a size a model can afford. GPT-2’s tokenizer has about 50,000 entries, the cl100k_base tokenizer used by GPT-4 has about 100,000, and o200k_base has about 200,000. That sounds large, but it is fixed and finite, unlike a list of every word.
Which pieces make the list is not decided by a linguist. The tokenizer learns them from a large sample of text, keeping the chunks that occur most often. That is why tokens rarely line up with dictionary syllables or prefixes. How that learning works is the subject of How tokenizers are built.
Bytes: the safety net
One question remains: what if even the smallest pieces are missing? There are well over 100,000 characters in Unicode, the standard that covers the world’s writing systems and emoji. Putting all of them in the vocabulary would waste space on characters that almost never appear.
The trick used by GPT-style tokenizers is to work on bytes instead. Computers store text as bytes using an encoding called UTF-8: a plain English letter is one byte, an accented letter like ü is two, and most emoji are four. There are only 256 possible byte values, so if all 256 are in the vocabulary, any text at all can be written down, one byte at a time if necessary.
You can see this in the German example. The giraffe emoji 🦒 is rare, so there is no single token for it. Instead it is split into its raw bytes: a token for a space plus the first two bytes F0 9F, then A6, then 92. None of those pieces is a character on its own, but together they rebuild the emoji exactly. Nothing is ever unknown; rare things just cost more tokens. (Why emoji and some languages cost more tokens looks at this in detail.)
Some tokenizers, such as the SentencePiece tokenizer used by the original LLaMA models, work on characters first and only drop down to bytes for characters they have never seen. The effect is the same: no text is ever unrepresentable.
What this means in practice
- Limits and prices are in tokens. Context windows and most API prices count tokens, not words. A common rule of thumb for English is that a token is about three quarters of a word, but it varies a lot with language and content.
- The same text can cost different amounts. Different models use different tokenizers, so the same sentence can be a different number of tokens in each.
- The model does not see letters.
·unbelievablyis one token with one ID. The model never directly sees the letters inside it, which is part of why spelling and letter-counting questions can trip models up.
Once text is a list of token IDs, the next step inside a model is to turn each ID into a list of meaningful numbers. That is the idea behind embeddings, and How a language model turns your words into an answer shows where both steps fit in the whole pipeline. To split your own text with several real tokenizers side by side, try the Tokenizer.
Further reading
- Byte-Pair Encoding tokenizationHugging Face LLM Course · chapter 6 covers tokenizers in depth
- Let’s build the GPT TokenizerAndrej Karpathy · video
- Byte pair encodingWikipedia
