← Learn · Foundations
Embeddings7 min read

What are embeddings?

Turning a word into a list of numbers so that similar meanings end up close together.

Computers are good with numbers and bad with meaning. Before a model can do anything useful with the word cat, it has to turn it into numbers. The interesting question is which numbers. An embedding is the answer almost every modern language system uses: a list of numbers for each word (or token), learned from text, arranged so that words with similar meanings get similar lists.

Why not just give each word a number?

The simplest idea is to number the vocabulary: the is 1, cat is 2, and so on. This is roughly what a tokenizer does (see Why models read tokens). But the numbers are arbitrary labels. In the 10,000-word list used below, cat is number 5,606 and dog is 3,037, but that says nothing about cats and dogs. A model would happily conclude that dog is “smaller” than cat, which is nonsense.

The classic fix is a one-hot vector: a list with one slot per vocabulary word, all zeros except a single 1 in that word’s slot. With a 10,000-word vocabulary, every word becomes 10,000 numbers. No more fake ordering, but a new problem appears: every pair of words is exactly as different as every other pair. cat is no closer to dog than to democracy. The representation carries no meaning at all.

Top: one-hot vectors, drawn with 16 slots instead of 10,000 (illustrative). Each word lights up its own slot, so no two words share anything. Bottom: the first 16 of 300 real GloVe numbers for the same words (teal positive, red negative). cat and dog have a similar pattern; car differs.

A short list of learned numbers

An embedding flips the approach. Each word gets a dense vector: a much shorter list (often 100 to 1,000 numbers, here 300) where every position holds some value. Nobody chooses these values by hand. They start random and are adjusted during training until words used in similar ways end up with similar lists.

If you treat each list of 300 numbers as coordinates, every word becomes a point in a 300-dimensional space. “Similar meaning” then becomes “nearby”. We can’t draw 300 dimensions, but we can squash a handful of words down to two and look.

catdoghorsecowlionbirdmousefishapplefruitbreadricecheesecoffeemilkfrancegermanyjapanparislondontokyoberlinredbluegreenyellowpurpletwofivetenkingqueenprincemanwomanboygirlcomputersoftwareinternetphonedigital

Nearest to apple among all 10,000 words

  1. microsoft0.64
  2. google0.61
  3. intel0.56
  4. software0.56
  5. ibm0.56
  6. computer0.53

Bold words are also on the map.

animals
food
places
colours
numbers
people
tech
Real GloVe data. 42 words placed in 2D so their distances match the 300-D distances as well as possible (a method called MDS, applied to these 42 words only, so the layout is approximate). The list under the map is exact: each word’s six nearest neighbours by cosine similarity among all 10,000 words. Try apple, rice and mouse.

The groups form on their own: the vectors were never told what an animal or a country is. The surprises are just as instructive.apple sits between food and tech, and its nearest neighbours are all companies, because the training text talks about Apple the company far more than the fruit.rice’s closest word is condoleezza, after the former US Secretary of State. mouse is pulled toward cartoons and computers. An embedding reflects how words are used in its training text, not what a dictionary says.

You shall know a word by the company it keeps

Why should “used in similar ways” produce “similar meaning”? This is the distributional hypothesis, summed up by the linguist J.R. Firth in 1957: “You shall know a word by the company it keeps.” Consider the gap in “I poured some ___ into my cup.” coffee, tea and milk all fit;bicycle does not. Words that fit the same gaps across millions of sentences are probably related, and that is something a computer can measure without understanding anything.

How the numbers are learned

Two famous methods from the 2010s show the idea clearly. Both read a large amount of text and produce one vector per word.

  • word2vec (Mikolov and colleagues at Google, 2013) plays a guessing game. Take a word from a sentence and try to predict the words around it (or the reverse). Each wrong guess slightly nudges the vectors involved. After billions of nudges, words that predict the same neighbours have been pushed toward similar vectors.
  • GloVe (Pennington, Socher and Manning at Stanford, 2014) starts by counting how often every pair of words appears near each other in the whole corpus. It then fits vectors so that comparing two words’ vectors predicts how often they co-occur.

Different recipes, same result: a table with one row of numbers per word. These are called static embeddings, because each word gets exactly one vector, whatever sentence it appears in.

One word, many meanings: contextual embeddings

Static vectors have an obvious weakness. bank can be the side of a river or a place that keeps money, but it only gets one vector. In this GloVe data its nearest neighbours are banks, banking, central, credit and financial. The river meaning is mostly drowned out.

She sat on the river bank and watched the water.

He moved his savings to another bank.

bank (static)bank + “river”bank + “savings”context mixes in
Illustrative. A static embedding gives “bank” one vector for both sentences. A contextual model starts from that same vector, then its layers mix in the surrounding words, so the two uses end up in different places.

Modern language models keep the lookup table as their first step: every token ID is swapped for a learned vector, exactly like above. But then the vectors pass through many layers of attention, where each token’s vector is updated using the tokens around it (see Attention and the transformer). By the later layers, “bank” in a sentence about rivers has a different vector from “bank” in a sentence about savings. These are contextual embeddings: the vector depends on the whole sentence, not just the word.

Measuring closeness: cosine similarity

To say two words are “near”, we need a way to measure it. The most common is cosine similarity. Picture each vector as an arrow from the origin. Cosine similarity is the cosine of the angle between two arrows: 1 when they point the same way, 0 when they are at right angles, and −1 when they point in opposite directions.

catdog · 0.72car · 0.28democracy · 0.15
Cosine similarity is about the angle between two arrows. Each arrow is drawn at the real angle between that word’s GloVe vector and cat’s (cat–dog 0.72, about 44°; cat–car 0.28, about 74°; cat–democracy 0.15, about 81°). Only the angles are meaningful here, not the lengths.

In practice you compute it by multiplying the two lists position by position, adding up the results (the dot product), and dividing by both arrows’ lengths. Dividing by length means only direction matters, which is useful because a vector’s length can reflect things like how often a word appears, rather than what it means. In the GloVe data, cat and dog score 0.72 and king and queen score 0.70, while cat and car, one letter apart, score only 0.28. The embedding knows nothing about spelling, only about usage.

Straight-line (Euclidean) distance works too, and when all vectors are scaled to length 1 it ranks neighbours in exactly the same order as cosine similarity.

Why this matters

Once meaning is a position, many tasks become geometry. Search engines find documents whose embeddings are close to your query’s. Recommendation systems suggest items near the ones you liked. Chatbots that look things up in your files (often called retrieval-augmented generation) embed every chunk of text and fetch the chunks nearest to the question. And inside every language model, embeddings are the form in which text first enters the network (see How a language model turns your words into an answer).

The space these vectors live in has more structure than just “near” and “far”: directions can carry meaning too. That is the subject of Latent space. How to draw 300 dimensions on a flat screen without fooling yourself is covered in Seeing high dimensions. To explore 10,000 real word vectors in 3D, open the Embedding explorer, and for a site-specific tour see How a word becomes 300 numbers.

Further reading