看懂你看到的一切。
语言模型如何读懂文字:从最初的一步,到它写出的下一个词。
不知道从哪里开始?从这里开始
基础知识
通用学习资料。每篇文章把一个概念讲透,就像一本好教材的一章。
Why models read tokens, not letters or words
The trade-off between letters, words and word pieces, and why every modern model settled on pieces, with bytes as a safety net.
How tokenizers are built: BPE, WordPiece and Unigram
Step by step through byte-pair encoding, and how WordPiece and Unigram make different choices.
Tokens beyond text: images, audio and video
How multimodal models cut pictures into patches and sound into frames so they can be read like words.
What are embeddings?
Turning a word into a list of numbers so that similar meanings end up close together.
Latent space: the hidden meaning in numbers
What the dimensions of an embedding capture, why directions carry meaning, and what we can't read directly.
Seeing high dimensions: PCA, t-SNE, UMAP and friends
How to squash hundreds of dimensions into 2 or 3, what each method keeps, and what it quietly throws away.
Attention and the transformer, explained from scratch
How each token looks at the others to update its meaning, and how stacking that idea builds a transformer.
From the last vector to the next word
Logits, softmax, temperature, top-k and top-p: how a model picks what to say next.
站点指南
快速了解每个 WordCanvas3D 工具展示了什么,并链接回基础知识。
What is a token, and why isn’t it a word?
How tokenizers cut text into reusable pieces, and why those pieces rarely line up with words.
配套工具:分词器Why emoji and other languages cost more tokens
UTF-8 bytes, byte-level BPE, and why the same greeting can take 4 tokens or 25.
配套工具:分词器How a word becomes 300 numbers
What word embeddings are, how GloVe, Word2Vec and FastText learn them, and how to measure similarity.
配套工具:词嵌入浏览器PCA vs UMAP: two ways to flatten meaning
Two ways to squeeze 300 dimensions into 3, what each one keeps, and how to read the result.
配套工具:词嵌入浏览器King − man + woman, explained
Word analogies as arrow arithmetic: what the Playground computes, why it works, and where it breaks.
配套工具:向量实验场
