← 学ぶ · サイトガイド
埋め込み4 分で読めます

How a word becomes 300 numbers

What word embeddings are, how GloVe, Word2Vec and FastText learn them, and how to measure similarity.

この記事は現在、英語のみです。サイトのほかの部分は翻訳されています。

A tokenizer turns text into IDs, but an ID is only a position in a list. Nothing about the number 9059 says “cat”. To work with meaning, a model needs a representation where similar words look similar. That representation is an embedding: a list of numbers, often a few hundred long, for every word or token. Such a list is also called a vector, and each position in it is a dimension.

This guide sticks to what the site’s Embedding page shows. For the idea itself, built up from scratch, read What are embeddings?

The Embedding page on this site shows three classic sets of word embeddings: GloVe, Word2Vec and FastText. In all three, every word is a list of exactly 300 numbers.

What 300 numbers look like

Here are the first 8 of the 300 numbers for four words, taken from the GloVe data this site uses.

king-0.07-0.37-0.21-0.880.06-0.260.34-0.23
queen-0.08-0.30-0.18-0.81-0.660.070.21-0.34
man-0.02-0.02-0.400.270.180.28-0.270.10
apple-0.170.130.34-0.71-0.21-0.77-0.440.26
First 8 of 300 GloVe values, rounded. Teal is positive, red is negative, and stronger colour means a larger value.

No single number means “royal” or “fruit”. The meaning is spread across all 300 at once. Still, even this tiny slice hints at a pattern:king and queen have a similar shape in several positions, while man and apple go their own way. What the dimensions do and don’t capture is the subject of Latent space.

Where the numbers come from

All three models are built on the same idea, often summed up by the linguist J.R. Firth: “You shall know a word by the company it keeps.” Words that appear in similar contexts tend to have similar meanings. The models read huge amounts of text and adjust each word’s numbers until words with similar neighbours end up with similar lists.

  • Word2Vec (Google, 2013) trains a small neural network on a simple guessing game: predict the words around a given word, or the word from the words around it. After training, the network’s internal weights for each word become its embedding.
  • GloVe (Stanford, 2014) first counts how often each pair of words appears near each other across the whole corpus. It then fits vectors so that the relationship between two words’ vectors reflects how often those words co-occur.
  • FastText (Facebook, 2016–17) extends Word2Vec by also learning vectors for chunks of characters inside words, so walk, walking and walked share parts. That helps with rare words and spelling variants.

The details differ in the data too. The GloVe vectors here are all lowercase. Word2Vec and FastText keep capital letters, so Paris and paris can be different entries, and the Word2Vec set even includes joined phrases like prime_minister.

Measuring similarity

If each word is a point in 300-dimensional space, “similar meaning” becomes “pointing in a similar direction”. The usual measure is cosine similarity: 1 means the two vectors point the same way, 0 means they are unrelated (at right angles).

PairSimilarity
cat · dog0.72
king · queen0.70
king · man0.47
cat · car0.28
king · apple0.27
Cosine similarity between word pairs, computed from the site’s GloVe vectors.

The numbers match intuition: cat is much closer to dog than to car, even though cat and car differ by one letter. The embedding knows nothing about spelling; it only knows usage.

What embeddings get wrong

Usage is also where the surprises come from. In this GloVe data, the nearest neighbours of apple are microsoft, google and intel. The training text talked about the company far more than the fruit.

Opposites can be close, too. good and bad have a similarity of about 0.72, because they appear in almost identical sentences: “the food was good”, “the food was bad”.

These models are also static: each word gets exactly one vector, no matter the sentence. The bank of a river and the bank that holds your money share a single point. Modern language models fix this by adjusting each token’s vector based on the words around it, layer by layer (Attention and the transformer shows how). But their very first step is the same as here: look up a vector for each token ID in a big table.

Why 300?

The length of the list is a design choice made before training. More numbers give the model more room to record fine distinctions, but they need more training data, more memory and more computation. For classic word embeddings like these, 300 became a common size. Across the 10,000 words each model ships with here, that is 3 million numbers per model.

Large language models use the same idea at a bigger scale, with token vectors that are often thousands of numbers long.

Seeing the space

Nobody can picture 300 dimensions. The Embedding page takes 1,000, 5,000 or 10,000 words from a model and squeezes their vectors down to 3D so you can fly through them. How that squeezing works, and what it hides, is covered in PCA vs UMAP.