PCA vs UMAP: two ways to flatten meaning
Two ways to squeeze 300 dimensions into 3, what each one keeps, and how to read the result.
Each word on the Embedding page is really a list of 300 numbers, a point in 300-dimensional space. Screens have two dimensions and our intuition handles three. To draw the words at all, the site has to reduce 300 numbers to 3. This is called dimensionality reduction, and the page offers two ways to do it: PCA and UMAP.
Neither is “correct”. Throwing away 297 dimensions always loses something. The two methods simply choose to keep different things. This guide covers the two options on the Embedding page; Seeing high dimensions explains the wider family, including t-SNE, in more depth.
PCA: find the widest directions
Principal Component Analysis looks for the direction in which the points are most spread out. That becomes the first axis. Then it finds the next most spread-out direction at right angles to the first, and so on. To get a 3D view, it keeps the top three directions and drops the rest.
Think of photographing a flat object like a book. You get the most information by shooting it face-on, where its shape is widest, not edge-on. PCA picks the camera angle that shows the most spread. Mathematically it is just a rotation followed by dropping axes (a projection, like a shadow cast onto a wall), so it has useful properties:
- It is deterministic: the same data always gives the same picture.
- It is linear, so straight-line relationships in the original space stay straight.
- Large distances are roughly kept: things far apart in the plot really are far apart in the data.
The weakness is how much it has to leave out. Word embeddings spread their information across many directions. For the 10,000 GloVe vectors this site ships, the top three directions capture only about 9% of the total spread; it takes around 54 directions to reach half. So a PCA view tends to look like one big cloud, with groups overlapping in the middle. Two words can land close together on screen without being similar at all.
UMAP: keep the neighbours together
UMAP (Uniform Manifold Approximation and Projection, 2018) takes a different approach. Instead of preserving overall spread, it tries to preserve neighbourhoods.
First it finds each word’s nearest neighbours in the full 300-dimensional space and builds a graph connecting them. Then it places the words in 3D and nudges them around until the 3D neighbours match the original ones as closely as possible: neighbours pull together, non-neighbours push apart.
PCA
UMAP
The result usually shows clear, tight clusters: numbers in one island, place names in another. That makes UMAP views easier to explore. But the clarity comes with rules for reading them:
- Distances between clusters mean little. Two islands far apart are not necessarily less related than two islands side by side.
- Cluster size means little. UMAP tends to even out density, so a large, loose group and a small, tight one can look alike.
- Runs differ. UMAP starts from a random layout, so running it again can rotate, flip or rearrange the islands. The views on this site are computed ahead of time, so they stay the same between visits.
Side by side
| PCA | UMAP | |
|---|---|---|
| Keeps | overall spread | local neighbours |
| Method | linear (rotation) | non-linear (graph) |
| Same result each run | yes | not guaranteed |
| Clusters look | overlapping | separated |
| Trust distances | large ones, roughly | small ones only |
Reading the Embedding page
On the Embedding page, the Dimensionality Reduction switch flips between the two. Try the same model and word count in both and compare.
- The colours come from cluster IDs stored with each dataset file. They are computed separately for each view, so a colour in PCA does not mean the same group as that colour in UMAP.
- The page itself warns that distance in the 3D view is not the original similarity score. For true similarity, the full 300 numbers are what count.
- Search for a word and look at its neighbours in both views. Neighbours that appear in both are a good sign the relationship is real, not an artefact of the projection.
A good habit is to use UMAP to find interesting groups, and PCA as a sanity check on the big picture.
