Learn

Understand what you’re looking at.

How language models read text, from the very first step to the next word they write.

Don’t know where to start? Start here

Foundations

General study material. Each article explains one idea properly, the way a good textbook chapter would.

  1. 01Tokens6 min read

    Why models read tokens, not letters or words

    The trade-off between letters, words and word pieces, and why every modern model settled on pieces, with bytes as a safety net.

  2. 02Tokens6 min read

    How tokenizers are built: BPE, WordPiece and Unigram

    Step by step through byte-pair encoding, and how WordPiece and Unigram make different choices.

  3. 03Tokens5 min read

    Tokens beyond text: images, audio and video

    How multimodal models cut pictures into patches and sound into frames so they can be read like words.

  4. 04Embeddings7 min read

    What are embeddings?

    Turning a word into a list of numbers so that similar meanings end up close together.

  5. 05Embeddings7 min read

    Latent space: the hidden meaning in numbers

    What the dimensions of an embedding capture, why directions carry meaning, and what we can't read directly.

  6. 06Embeddings7 min read

    Seeing high dimensions: PCA, t-SNE, UMAP and friends

    How to squash hundreds of dimensions into 2 or 3, what each method keeps, and what it quietly throws away.

  7. 07Transformers7 min read

    Attention and the transformer, explained from scratch

    How each token looks at the others to update its meaning, and how stacking that idea builds a transformer.

  8. 08Transformers5 min read

    From the last vector to the next word

    Logits, softmax, temperature, top-k and top-p: how a model picks what to say next.

Site guides

Quick tours of what each WordCanvas3D tool shows, with links back to the foundations.

  1. Tokens4 min read

    What is a token, and why isn’t it a word?

    How tokenizers cut text into reusable pieces, and why those pieces rarely line up with words.

    Pairs with: Tokenizer
  2. Tokens4 min read

    Why emoji and other languages cost more tokens

    UTF-8 bytes, byte-level BPE, and why the same greeting can take 4 tokens or 25.

    Pairs with: Tokenizer
  3. Embeddings4 min read

    How a word becomes 300 numbers

    What word embeddings are, how GloVe, Word2Vec and FastText learn them, and how to measure similarity.

    Pairs with: Embedding explorer
  4. Embeddings4 min read

    PCA vs UMAP: two ways to flatten meaning

    Two ways to squeeze 300 dimensions into 3, what each one keeps, and how to read the result.

    Pairs with: Embedding explorer
  5. Vectors4 min read

    King − man + woman, explained

    Word analogies as arrow arithmetic: what the Playground computes, why it works, and where it breaks.

    Pairs with: Vector Playground
Back to top