Aprender

Entiende lo que estás viendo.

Cómo los modelos de lenguaje leen el texto, desde el primer paso hasta la siguiente palabra que escriben.

¿No sabes por dónde empezar? Empieza aquí

Fundamentos

Material de estudio general. Cada artículo explica bien una idea, como lo haría un buen capítulo de libro de texto.

  1. 01Tokens6 min de lectura

    Why models read tokens, not letters or words

    The trade-off between letters, words and word pieces, and why every modern model settled on pieces, with bytes as a safety net.

  2. 02Tokens6 min de lectura

    How tokenizers are built: BPE, WordPiece and Unigram

    Step by step through byte-pair encoding, and how WordPiece and Unigram make different choices.

  3. 03Tokens5 min de lectura

    Tokens beyond text: images, audio and video

    How multimodal models cut pictures into patches and sound into frames so they can be read like words.

  4. 04Embeddings7 min de lectura

    What are embeddings?

    Turning a word into a list of numbers so that similar meanings end up close together.

  5. 05Embeddings7 min de lectura

    Latent space: the hidden meaning in numbers

    What the dimensions of an embedding capture, why directions carry meaning, and what we can't read directly.

  6. 06Embeddings7 min de lectura

    Seeing high dimensions: PCA, t-SNE, UMAP and friends

    How to squash hundreds of dimensions into 2 or 3, what each method keeps, and what it quietly throws away.

  7. 07Transformers7 min de lectura

    Attention and the transformer, explained from scratch

    How each token looks at the others to update its meaning, and how stacking that idea builds a transformer.

  8. 08Transformers5 min de lectura

    From the last vector to the next word

    Logits, softmax, temperature, top-k and top-p: how a model picks what to say next.

Guías del sitio

Recorridos rápidos de lo que muestra cada herramienta de WordCanvas3D, con enlaces a los fundamentos.

  1. Tokens4 min de lectura

    What is a token, and why isn’t it a word?

    How tokenizers cut text into reusable pieces, and why those pieces rarely line up with words.

    Va con: Tokenizador
  2. Tokens4 min de lectura

    Why emoji and other languages cost more tokens

    UTF-8 bytes, byte-level BPE, and why the same greeting can take 4 tokens or 25.

    Va con: Tokenizador
  3. Embeddings4 min de lectura

    How a word becomes 300 numbers

    What word embeddings are, how GloVe, Word2Vec and FastText learn them, and how to measure similarity.

    Va con: Explorador de embeddings
  4. Embeddings4 min de lectura

    PCA vs UMAP: two ways to flatten meaning

    Two ways to squeeze 300 dimensions into 3, what each one keeps, and how to read the result.

    Va con: Explorador de embeddings
  5. Vectores4 min de lectura

    King − man + woman, explained

    Word analogies as arrow arithmetic: what the Playground computes, why it works, and where it breaks.

    Va con: Laboratorio de vectores
Volver arriba