Why emoji and other languages cost more tokens
UTF-8 bytes, byte-level BPE, and why the same greeting can take 4 tokens or 25.
本文暂时只有英文版,网站的其他部分已翻译。
Type “Hello, world.” into the Tokenizer and you get 4 tokens. Type the same greeting in Hindi and, depending on the tokenizer, you can get anywhere from 6 to 25. Emoji behave similarly: a single 🎉 can take two, three or more tokens. None of this is random. It comes from how text is stored as bytes, and from what the tokenizer saw during training.
Everything starts as bytes
Computers store text using an encoding, a fixed rule for turning each character into bytes (a byte is a small number from 0 to 255, the basic unit of computer memory). Almost everything today uses UTF-8, which gives each character between one and four bytes:
- Basic Latin letters, digits and punctuation: 1 byte each.
- Arabic, Hebrew, Greek and Cyrillic letters: 2 bytes each.
- Devanagari (used for Hindi), Chinese and Japanese characters: 3 bytes each.
- Most emoji: 4 bytes each.
You can watch this in the Tokenizer: under the input it shows the number of code points (the numbered characters of the Unicode standard; roughly, characters) and UTF-8 bytes. “Hello, world.” is 13 characters and 13 bytes. The Hindi “नमस्ते दुनिया।” is 14 code points but 40 bytes.
Merges follow the training data
The GPT tokenizers are byte-level BPE tokenizers. They start from single bytes and learn merges from a large pile of training text, keeping the pairs that appear most often. (The first guide, What is a token, gives the short version;How tokenizers are built walks through the merges step by step.)
If the training text is mostly English, most merges end up being English words and word parts. Other scripts get fewer merges, so their text stays split into smaller pieces. In the worst case a single character is split across tokens, because its bytes were never merged together.
Two things stand out. First, English costs the same 4 tokens everywhere, because every one of these vocabularies learned Hello, , and ·world long ago. Second, the newer o200k_base is far better at the other scripts. Its vocabulary is about twice the size of cl100k_base’s, and its extra entries include many more pieces of non-English text. 世界 (“world” in Japanese) is one token in o200k_base, while cl100k_base needs three, splitting 世 into raw bytes.
Emoji: several bytes, sometimes several characters
The 🎉 emoji is one character stored as four bytes: F0 9F 8E 89. A tokenizer that has not merged those bytes into one piece has to spend several tokens on it. The Tokenizer shows pieces that are only part of a character as hex bytes, so you can see exactly where the split happens.
Popular emoji like 👍 and ✅ are single tokens in o200k_base, while in cl100k_base they still take two or three.
Some emoji are not even one character. 🛠️ is the hammer-and-wrench symbol plus an invisible variation selector that asks for the colourful emoji style. That makes 4 tokens in both GPT-4 era tokenizers. Family emoji such as 👨👩👧 are several people glued together with invisible “zero width joiner” characters: 5 code points, 18 bytes, and 8 tokens in o200k_base, for what looks like one symbol.
Why the count matters
- Cost. Many AI APIs charge per token, so the same message can cost more in one language than another.
- Context. A model’s context window is a token budget. Text that needs more tokens fills it sooner.
- Speed. Models generate one token at a time, so a reply that needs more tokens takes longer to write.
This is also why tokenizer changes matter. Moving from cl100k_base to o200k_base cut the Hindi greeting above from 15 tokens to 6, with no change to the text.
Practical takeaways
- Count with the right tokenizer. The same text gives different counts in different vocabularies, so a count from one model does not carry over to another.
- Decoration is not free. Emoji, box-drawing characters and fancy symbols in a prompt can each cost several tokens while adding little meaning.
- Nothing is ever “unknown”. Because byte-level tokenizers can always fall back to single bytes, any text can be encoded. Rare text is not rejected, it is just expensive.
Try it
The Tokenizer has two presets made for this article: World scripts and Emoji. Load one, then switch between tokenizers and watch the token count change. Turn on the LLaMA tokenizer too: it uses roughly one token per character for these scripts, and to raw bytes for most emoji.
