← Aprender · Fundamentos
Tokens5 min de lectura

Tokens beyond text: images, audio and video

How multimodal models cut pictures into patches and sound into frames so they can be read like words.

Por ahora, este artículo solo está disponible en inglés. El resto del sitio está traducido.

A transformer, the design behind modern language models, does not really care what its tokens are. All it ever receives is a sequence of vectors (lists of numbers), one per token, and its job is to let those vectors look at each other and update themselves. For text, each vector comes from a word piece. But nothing stops us from making vectors out of other things.

That is how models that see pictures and hear sound work. The trick is always the same: cut the input into small pieces, turn each piece into a vector, and line the vectors up as a sequence. This article shows how that is done for images, audio and video, and how the result is mixed with text.

Images: an image is worth 16×16 words

A digital image is a grid of pixels, and each pixel is three numbers: how red, green and blue it is. A single pixel is far too small to mean anything, just like a single letter. Treating every pixel as a token would also produce enormous sequences: a modest 224 × 224 photo has over 50,000 pixels.

The Vision Transformer (ViT), published by Google researchers in 2020 under the title “An Image is Worth 16x16 Words”, made the idea simple. Cut the image into a grid of small squares called patches, typically 16 × 16 pixels each, and treat every patch as one token. The patches are read in order, left to right and top to bottom, like words on a page.

0 of 16 patches in the sequence
image, cut into a 4 × 4 grid
one input sequence for the model
  1. What
  2. ·is
  3. ·in
  4. ·this
  5. ·picture
  6. ?
  7. <image>
  8. 1
  9. 2
  10. 3
  11. 4
  12. 5
  13. 6
  14. 7
  15. 8
  16. 9
  17. 10
  18. 11
  19. 12
  20. 13
  21. 14
  22. 15
  23. 16
  24. </image>
Illustrative: an image cut into a 4 × 4 grid of patches, which join the same sequence as the text tokens of a question. Press Play or step through one patch at a time. Real models use many more patches, and marker tokens like <image> vary from model to model.

From a patch to a vector

A patch is still just pixels. To turn it into a token vector, the model does two simple things.

  • Flatten it. A 16 × 16 patch with three colour values per pixel is 16 × 16 × 3 = 768 numbers. Lay them out in one long list.
  • Project it. Multiply that list by a learned matrix to get a vector of the size the transformer expects. This plays the same role as the lookup table that turns a text token ID into an embedding. (In the base-sized ViT, the output also happens to be 768 numbers long.)

Finally, a position embedding is added to each vector, so the model knows where in the image the patch came from; without it, a shuffled image would look the same. From here on, the patch vectors go through ordinary attention layers, where every patch can look at every other patch.

Patch size is a trade-off, just like vocabulary size for text. Smaller patches keep more detail but produce many more tokens, and the cost of attention grows quickly with sequence length.

image size
patch size
patches per side
224 ÷ 16 = 14
image tokens
14 × 14 = 196
numbers in one patch
16 × 16 × 3 colours = 768
How many tokens an image becomes. Halving the patch size, or doubling the image size, gives four times as many tokens. The arithmetic is exact; 224 × 224 pixels with 16 × 16 patches (196 tokens) is the standard ViT setup.

Continuous patches or discrete codes?

There is one real difference from text. A text token is one entry from a fixed vocabulary, so the model can predict it by picking from a list. A patch vector is continuous: its numbers can be anything, and there is no list. For understanding images that is fine; most vision encoders, including ViT, work this way.

For generating images one token at a time, some models first turn images into discrete tokens. A separately trained network, such as a VQ-VAE (vector-quantised variational autoencoder), learns a codebook: a fixed list of typical little image pieces. Each region of an image is replaced by the ID of its closest codebook entry, so an image becomes a grid of IDs, exactly like text. The original DALL·E, for example, turned each 256 × 256 image into a 32 × 32 grid of 1,024 image tokens, each chosen from a codebook of 8,192 entries, and placed them after up to 256 text tokens in one sequence.

Continuous patch embeddingsDiscrete image tokens
A token isA vector computed from a patch’s pixelsAn ID from a learned codebook
VocabularyNone: any vector is possibleFixed, e.g. 8,192 codes in DALL·E
Good forUnderstanding: classifying, describing, answering questionsGenerating images token by token, like text
ExampleViT: 16 × 16 pixel patchesDALL·E: 32 × 32 grid of codes
The two ways of turning a picture into tokens. Examples are from the ViT, VQ-VAE and DALL·E papers.

Audio: slices of sound

Sound is a wave: a long list of air-pressure measurements, often 16,000 or more per second. That is far too many to use one per token, and single measurements mean nothing on their own.

So speech models usually start by computing a spectrogram, a picture of which frequencies are loud at each moment. Time runs along one axis and pitch along the other, and each short slice of time becomes a column of numbers. OpenAI’s Whisper speech recogniser, for instance, takes 30-second clips, computes an 80-channel spectrogram with a new slice every 10 milliseconds (3,000 slices), and two small convolution layers then halve that to 1,500 vectors before the transformer reads them.

sound wavespectrogram: pitch × timeone vector per slice… a sequence of audio tokensslices every few milliseconds
Illustrative: from a sound wave to a spectrogram to a sequence of audio vectors. One highlighted slice of time becomes one token. The spectrogram colours here are drawn by hand, not computed from real audio.

Audio has its discrete version too. Neural audio codecs, such as Meta’s EnCodec, compress sound into a stream of codebook IDs, several per slice of time. Those IDs can be predicted one after another like text, which is how some speech and music generators produce audio.

Video: patches in space and time

A video is a stack of images. The simplest approach tokenizes each frame into patches, but that multiplies the token count by the number of frames. Video transformers such as ViViT instead cut the video into tubelets: small boxes that span a patch of the picture and a few consecutive frames. Each tubelet is flattened and projected into one vector, so a single token captures a little bit of motion as well as appearance. Even so, video is expensive: a few seconds can produce thousands of tokens, which is why video models often sample only some of the frames.

Putting it all in one sequence

A multimodal model is often built from parts. A common recipe, used by the open-source LLaVA model, is:

  • a pre-trained vision encoder (in LLaVA’s case, CLIP’s ViT) turns the image into patch vectors;
  • a small learned projection maps each of those vectors into the same space as the language model’s word embeddings, so they have the right length and “speak the same language”;
  • the resulting image tokens are placed in the sequence next to the text tokens, and the language model reads them all together.

Inside the transformer there is no special treatment. A text token can attend to an image patch exactly as it attends to another word, which is how a question like “What colour is the sun?” finds the patch that contains the sun.

The pieces also cost the same thing: space in the context window. An image or a minute of audio can use as many tokens as several paragraphs of text, which is why the same limits and prices apply. For how text tokens are chosen, see Why models read tokens; for what happens to all these vectors once they are in the sequence, see How a language model turns your words into an answer.

Further reading