学ぶ

目の前のものを理解しよう。

言語モデルが文章を読むしくみを、最初のステップからモデルが書く次の単語まで。

どこから始めればいいかわからない?ここから始めよう

基礎

一般的な学習用の教材です。良い教科書の 1 章のように、各記事で 1 つの考え方をしっかり説明します。

  1. 01トークン6 分で読めます

    Why models read tokens, not letters or words

    The trade-off between letters, words and word pieces, and why every modern model settled on pieces, with bytes as a safety net.

  2. 02トークン6 分で読めます

    How tokenizers are built: BPE, WordPiece and Unigram

    Step by step through byte-pair encoding, and how WordPiece and Unigram make different choices.

  3. 03トークン5 分で読めます

    Tokens beyond text: images, audio and video

    How multimodal models cut pictures into patches and sound into frames so they can be read like words.

  4. 04埋め込み7 分で読めます

    What are embeddings?

    Turning a word into a list of numbers so that similar meanings end up close together.

  5. 05埋め込み7 分で読めます

    Latent space: the hidden meaning in numbers

    What the dimensions of an embedding capture, why directions carry meaning, and what we can't read directly.

  6. 06埋め込み7 分で読めます

    Seeing high dimensions: PCA, t-SNE, UMAP and friends

    How to squash hundreds of dimensions into 2 or 3, what each method keeps, and what it quietly throws away.

  7. 07Transformer7 分で読めます

    Attention and the transformer, explained from scratch

    How each token looks at the others to update its meaning, and how stacking that idea builds a transformer.

  8. 08Transformer5 分で読めます

    From the last vector to the next word

    Logits, softmax, temperature, top-k and top-p: how a model picks what to say next.

サイトガイド

WordCanvas3D の各ツールで何が見えるのかを手短に紹介し、基礎の記事へのリンクも付けています。

  1. トークン4 分で読めます

    What is a token, and why isn’t it a word?

    How tokenizers cut text into reusable pieces, and why those pieces rarely line up with words.

    関連ツール: トークナイザー
  2. トークン4 分で読めます

    Why emoji and other languages cost more tokens

    UTF-8 bytes, byte-level BPE, and why the same greeting can take 4 tokens or 25.

    関連ツール: トークナイザー
  3. 埋め込み4 分で読めます

    How a word becomes 300 numbers

    What word embeddings are, how GloVe, Word2Vec and FastText learn them, and how to measure similarity.

    関連ツール: 埋め込みエクスプローラー
  4. 埋め込み4 分で読めます

    PCA vs UMAP: two ways to flatten meaning

    Two ways to squeeze 300 dimensions into 3, what each one keeps, and how to read the result.

    関連ツール: 埋め込みエクスプローラー
  5. ベクトル4 分で読めます

    King − man + woman, explained

    Word analogies as arrow arithmetic: what the Playground computes, why it works, and where it breaks.

    関連ツール: ベクトル・プレイグラウンド
トップに戻る