What are embeddings?
Turning a word into a list of numbers so that similar meanings end up close together.
Computers are good with numbers and bad with meaning. Before a model can do anything useful with the word cat, it has to turn it into numbers. The interesting question is which numbers. An embedding is the answer almost every modern language system uses: a list of numbers for each word (or token), learned from text, arranged so that words with similar meanings get similar lists.
Why not just give each word a number?
The simplest idea is to number the vocabulary: the is 1, cat is 2, and so on. This is roughly what a tokenizer does (see Why models read tokens). But the numbers are arbitrary labels. In the 10,000-word list used below, cat is number 5,606 and dog is 3,037, but that says nothing about cats and dogs. A model would happily conclude that dog is “smaller” than cat, which is nonsense.
The classic fix is a one-hot vector: a list with one slot per vocabulary word, all zeros except a single 1 in that word’s slot. With a 10,000-word vocabulary, every word becomes 10,000 numbers. No more fake ordering, but a new problem appears: every pair of words is exactly as different as every other pair. cat is no closer to dog than to democracy. The representation carries no meaning at all.
A short list of learned numbers
An embedding flips the approach. Each word gets a dense vector: a much shorter list (often 100 to 1,000 numbers, here 300) where every position holds some value. Nobody chooses these values by hand. They start random and are adjusted during training until words used in similar ways end up with similar lists.
If you treat each list of 300 numbers as coordinates, every word becomes a point in a 300-dimensional space. “Similar meaning” then becomes “nearby”. We can’t draw 300 dimensions, but we can squash a handful of words down to two and look.
The groups form on their own: the vectors were never told what an animal or a country is. The surprises are just as instructive.apple sits between food and tech, and its nearest neighbours are all companies, because the training text talks about Apple the company far more than the fruit.rice’s closest word is condoleezza, after the former US Secretary of State. mouse is pulled toward cartoons and computers. An embedding reflects how words are used in its training text, not what a dictionary says.
You shall know a word by the company it keeps
Why should “used in similar ways” produce “similar meaning”? This is the distributional hypothesis, summed up by the linguist J.R. Firth in 1957: “You shall know a word by the company it keeps.” Consider the gap in “I poured some ___ into my cup.” coffee, tea and milk all fit;bicycle does not. Words that fit the same gaps across millions of sentences are probably related, and that is something a computer can measure without understanding anything.
How the numbers are learned
Two famous methods from the 2010s show the idea clearly. Both read a large amount of text and produce one vector per word.
- word2vec (Mikolov and colleagues at Google, 2013) plays a guessing game. Take a word from a sentence and try to predict the words around it (or the reverse). Each wrong guess slightly nudges the vectors involved. After billions of nudges, words that predict the same neighbours have been pushed toward similar vectors.
- GloVe (Pennington, Socher and Manning at Stanford, 2014) starts by counting how often every pair of words appears near each other in the whole corpus. It then fits vectors so that comparing two words’ vectors predicts how often they co-occur.
Different recipes, same result: a table with one row of numbers per word. These are called static embeddings, because each word gets exactly one vector, whatever sentence it appears in.
One word, many meanings: contextual embeddings
Static vectors have an obvious weakness. bank can be the side of a river or a place that keeps money, but it only gets one vector. In this GloVe data its nearest neighbours are banks, banking, central, credit and financial. The river meaning is mostly drowned out.
Modern language models keep the lookup table as their first step: every token ID is swapped for a learned vector, exactly like above. But then the vectors pass through many layers of attention, where each token’s vector is updated using the tokens around it (see Attention and the transformer). By the later layers, “bank” in a sentence about rivers has a different vector from “bank” in a sentence about savings. These are contextual embeddings: the vector depends on the whole sentence, not just the word.
Measuring closeness: cosine similarity
To say two words are “near”, we need a way to measure it. The most common is cosine similarity. Picture each vector as an arrow from the origin. Cosine similarity is the cosine of the angle between two arrows: 1 when they point the same way, 0 when they are at right angles, and −1 when they point in opposite directions.
In practice you compute it by multiplying the two lists position by position, adding up the results (the dot product), and dividing by both arrows’ lengths. Dividing by length means only direction matters, which is useful because a vector’s length can reflect things like how often a word appears, rather than what it means. In the GloVe data, cat and dog score 0.72 and king and queen score 0.70, while cat and car, one letter apart, score only 0.28. The embedding knows nothing about spelling, only about usage.
Straight-line (Euclidean) distance works too, and when all vectors are scaled to length 1 it ranks neighbours in exactly the same order as cosine similarity.
Why this matters
Once meaning is a position, many tasks become geometry. Search engines find documents whose embeddings are close to your query’s. Recommendation systems suggest items near the ones you liked. Chatbots that look things up in your files (often called retrieval-augmented generation) embed every chunk of text and fetch the chunks nearest to the question. And inside every language model, embeddings are the form in which text first enters the network (see How a language model turns your words into an answer).
The space these vectors live in has more structure than just “near” and “far”: directions can carry meaning too. That is the subject of Latent space. How to draw 300 dimensions on a flat screen without fooling yourself is covered in Seeing high dimensions. To explore 10,000 real word vectors in 3D, open the Embedding explorer, and for a site-specific tour see How a word becomes 300 numbers.
Further reading
- The Illustrated Word2vecJay Alammar · a visual walk through how word2vec is trained
- Transformers, the tech behind LLMs3Blue1Brown · includes a beautiful section on word embeddings
- Efficient Estimation of Word Representations in Vector SpaceMikolov et al., 2013 · the original word2vec paper
- GloVe: Global Vectors for Word RepresentationStanford NLP · project page with paper and downloads
