What is a token, and why isn’t it a word?
How tokenizers cut text into reusable pieces, and why those pieces rarely line up with words.
When you type a sentence into a chatbot, the model never sees your letters, and it never sees your words either. It sees a list of whole numbers. The step that turns text into those numbers is called tokenization, and each piece of text that gets its own number is a token.
Here is a short sentence split by the o200k_base tokenizer, the one used by GPT-4o and the default in this site’s Tokenizer. Each chip is one token, and the number under it is that token’s ID: its position in the tokenizer’s list of known pieces.
So far it looks like “one word, one token”. Two details already break that idea. The full stop is its own token. And the space in front of each word is glued onto the word: the model’s vocabulary has a separate entry for ·cat (with a space) and cat (without one).
The vocabulary is a fixed list
A tokenizer comes with a vocabulary: a list of text pieces, each paired with an ID. The name o200k_base hints at its size, roughly 200,000 entries. The older cl100k_base (GPT-4 and GPT-3.5) has about 100,000, and GPT-2’s has about 50,000.
Because the list is fixed, the tokenizer has to express any input with pieces it already has. Capitalization, spacing and spelling all matter. In o200k_base, cat is ID 8837, ·cat is 9059, and Cat is 23546. To the model these are three unrelated numbers. It only learns that they mean nearly the same thing by seeing them used in similar ways during training.
Why not just use words?
A word-level vocabulary sounds simpler, but it runs into trouble fast. There is no end to words: names, typos, slang, product names, code identifiers, words from other languages. Any word missing from the list would have to be replaced with a generic “unknown” marker, and its meaning would be lost.
The opposite extreme, one token per character, never meets an unknown word, but it makes every text very long. Models do more work for longer inputs, and a single letter carries almost no meaning on its own.
Modern tokenizers sit in between. They use subword pieces (smaller than a word, usually bigger than a letter): common words get a single token, and rarer words are built from a few smaller, reusable parts. The full trade-off is explained in Why models read tokens, not letters or words.
That last row is worth a second look. In the middle of a sentence, where the word follows a space, it is common enough to earn its own token. At the very start of a text, with no space before it, the same word is rare, so it gets built from parts.
How the pieces are chosen
The GPT tokenizers use a method called byte pair encoding (BPE). Training starts from the smallest possible pieces, individual bytes of text. It then scans a large amount of text, finds the pair of neighbouring pieces that appears most often, and merges that pair into a new vocabulary entry. It repeats this until the vocabulary reaches its target size.
Early merges capture very common letter pairs such as in or he. Later merges build whole frequent words and word endings like ization. That is why tokenization comes out as token + ization in all three GPT tokenizers. The pieces are statistical, not grammatical: they follow what was frequent in the training text, not where a dictionary would split a word. For a worked, merge-by-merge example, see How tokenizers are built.
The LLaMA options in the Tokenizer use a different tool, SentencePiece. You will notice two differences there. Its vocabulary writes spaces as ▁ (click a token to see this as its “vocabulary piece”), and every text starts with a special <s> token that marks the beginning of a sequence.
Side effects you can see
Once you know the model reads tokens, a few familiar quirks make more sense.
- Spelling questions are awkward. In
o200k_base,strawberryis three tokens:st,raw,berry. The model never directly sees ten separate letters, which is one reason letter-counting questions trip models up. - Numbers are chunked. The GPT tokenizers split
1234567into123,456,7. LLaMA’s tokenizer splits it into one token per digit. Arithmetic on chunks like these is harder than it looks. - Length limits are counted in tokens. A model’s context window (the most text it can read at once) and many API prices are measured in tokens, not words or characters.
What happens next
A token ID on its own is just a position in a list. ID 9059 is not “more” than ID 402 in any meaningful way. The next step inside a model is to swap each ID for a long list of numbers, called an embedding, that does carry meaning. That is the subject of How a word becomes 300 numbers.
