← 学习 · 基础知识
嵌入7 分钟阅读

Seeing high dimensions: PCA, t-SNE, UMAP and friends

How to squash hundreds of dimensions into 2 or 3, what each method keeps, and what it quietly throws away.

本文暂时只有英文版,网站的其他部分已翻译。

Word embeddings have hundreds of numbers per word; the vectors inside large language models have thousands. Our eyes handle two dimensions on a screen and three at most. Dimensionality reduction is the family of methods that squash many dimensions down to two or three so we can look. It is one of the most useful tools for understanding embeddings, and one of the easiest to misread. This article explains the main methods, what each one keeps, and what each one quietly throws away.

Why any picture must lie a little

Going from 300 numbers to 2 means discarding information. There is no way around it: a flat picture simply has fewer places to put things. A photo of a room loses depth; a world map stretches Greenland. Every method makes a choice about which relationships to keep.

light2DABA and B
Illustrative. Projecting is like casting a shadow. Points A and B are far apart in 3D, but they lie on the same line of sight, so their shadows land on the same spot. Every projection to fewer dimensions loses some differences like this one.

So the useful question is never “which method is correct?” but “what did this method try to keep, and what did it give up?”

A test shape: the swiss roll

To compare methods, it helps to use data where we know the right answer. A classic test is the swiss roll: a flat rectangular sheet, rolled up into a spiral in 3D. The points’ real structure is two-dimensional (a position along the sheet and a position across it), but the roll hides that. A method that truly understood the data would unroll it. Colour follows position along the sheet, so a good 2D picture keeps the colours in order.

Original data · 3D

PCA · 2D

PCA. PCA looks for the directions of greatest spread. The roll is widest across its spiral, so PCA shows it end-on: a spiral, not a sheet. Colours that are far apart along the sheet sit next to each other on neighbouring turns.

Click any point: it and its 12 nearest neighbours in 3D are outlined in white in both views. Do they stay together in 2D?

Illustrative data (700 points). PCA, t-SNE, Isomap and LDA are real scikit-learn runs; the UMAP view comes from a small re-implementation of the UMAP algorithm, so it is close to, but not identical to, the official library. Turn the 3D roll with the slider, then switch between methods to see each one’s 2D result. The points glide from one layout to the next so you can follow them.

Keep this picture in mind as we go through the methods one by one.

PCA: the widest directions

Principal component analysis (PCA) finds the direction in which the data is most spread out and calls it the first principal component. The second is the most spread-out direction at right angles to the first, and so on. Keeping the first two and dropping the rest gives a 2D picture. Geometrically this is a rotation followed by a shadow: PCA picks the camera angle that shows the most spread.

PCA is linear: it can only rotate and flatten, never bend. That makes it fast, predictable (the same data always gives the same picture) and honest about large distances. It also means it cannot unroll the swiss roll; it just photographs it from the side where it looks biggest. For word embeddings, where meaning is spread over many directions, the first two components often capture only a small share of the total spread, so PCA pictures of words tend to look like one big overlapping cloud.

t-SNE: keep the neighbours together

t-SNE (t-distributed stochastic neighbour embedding, 2008) gives up on large distances and focuses on neighbourhoods. For each point it asks, in the original space, “which other points are my close neighbours, and how close?” It then starts from a random or rough 2D layout and moves the points around, step by step, until each point’s 2D neighbours match its original neighbours as well as possible.

That makes t-SNE very good at revealing clusters. It also produces three well-known illusions:

  • Cluster sizes mean nothing. t-SNE adapts to local density, so a tight group and a loose group come out looking about the same size.
  • Distances between clusters mean little. Two clusters far apart on the plot may or may not be far apart in the data.
  • Settings change the picture. The main knob, perplexity, is roughly how many neighbours each point pays attention to. Typical values are between 5 and 50.

PCA (for reference)

t-SNE, perplexity 2

t-SNE, perplexity 5

t-SNE, perplexity 30

Illustrative data, real t-SNE runs. Three groups in 10 dimensions: red is very tight, teal is six times wider and close to red, yellow is far from both. PCA (a straight projection) keeps those differences in size and distance. t-SNE makes the groups look similar in size, and how far apart they appear depends on the perplexity setting.

t-SNE is also random: run it twice with different seeds and the picture can change, even though the neighbourhoods are similar.

UMAP: similar goal, faster, a bit more global

UMAP (uniform manifold approximation and projection, 2018) follows the same basic plan: build a graph connecting each point to its nearest neighbours, then find a 2D layout that keeps that graph intact. Its maths is different, and in practice it is much faster on large datasets, so it has become a common default for embeddings.

UMAP tends to keep somewhat more of the large-scale layout than t-SNE: groups that are related often end up nearer each other. But it is still a neighbourhood method, so the same caution applies: distances between clusters and the size of each cluster are not reliable. Its two main settings are the number of neighbours (small values focus on fine detail, large values on the big picture) and a minimum distance that controls how tightly points are packed.

Two more: Isomap and LDA

Isomap (2000) also builds a nearest-neighbour graph, but then measures the distance between two points as the shortest route through that graph, like walking along the surface instead of tunnelling through the air. These are called geodesic distances. It then lays points out so those distances are kept. On a clean shape like the swiss roll it works beautifully; on noisy data a few wrong links (“short circuits” between layers) can ruin it.

LDA here means linear discriminant analysis. Unlike everything above, it is supervised: you give it a label for each point (say, the topic of each word) and it finds the flat projection that best separates those labels. It is useful when you already know the groups and want to see how separable they are. Watch out for the name clash: in text analysis, “LDA” more often means latent Dirichlet allocation, an unrelated method for finding topics in documents.

At a glance

MethodTypeKeepsDon’t trustSpeed
PCALinearOverall spread; big distancesCurved shapes; small groups hidden in the crowdFast
t-SNENon-linearEach point’s close neighboursCluster sizes; distances between clustersSlower on big data
UMAPNon-linearClose neighbours, plus some large-scale layoutExact densities and distancesFast
IsomapNon-linearDistances measured along the data’s surfaceBreaks if neighbours link across gaps; noisy dataMedium
LDALinear, uses labelsSeparation between known classesEverything the labels don’t coverFast
What each method tries to keep and what it is known to distort. “Linear” methods can only rotate and flatten; non-linear ones can bend and stretch.

How to read these plots

  • Trust neighbours more than distances. Points drawn right next to each other are usually close in the original space. Points far apart may or may not be.
  • Ignore cluster sizes and gaps in t-SNE and UMAP. A big blob is not necessarily a varied group, and a wide gap is not necessarily a big difference.
  • Check the original space. Before believing a pattern, confirm it with real distances or similarities in the full vectors.
  • Try more than one setting and seed. A pattern that survives different perplexities, neighbour counts and random seeds is more likely to be real.
  • Random data can look structured. Neighbourhood methods can produce clumps even from noise, so be wary of clusters you can’t explain.

For why the structure is there in the first place, see Latent space and What are embeddings?. The site guide PCA vs UMAP shows how these two methods are used in the Embedding explorer, where you can switch between them on 10,000 real words.

Further reading