Back to archive

Neural Collapse

Neural Collapse

You train a model to recognize cats, dogs, and birds in photos. Inside it, each photo gets a list of numbers describing features. After successful training, do these descriptions remain chaotic or form a pattern?

Neural Collapse is a phenomenon observed late in training many classifiers: descriptions of examples in the same class approach a shared center, while class centers form a regular arrangement. A classifier is a model assigning an input to a category.

Imagine a plot with three clusters of points. Initially “cat” points are scattered. Later the cluster narrows, and the centers of cats, dogs, and birds become more symmetric. This illustrates geometry, rather than showing an actual plot of a particular network. In a space with more dimensions, researchers also observe final-layer weights aligning with these centers.

Papyan, Han, and Donoho, §2, describe four related properties of this phenomenon. “Collapse” does not mean a model failure or loss of all abilities. Results concern specified classification conditions; they guarantee neither similar behavior in any language model nor automatic robustness to new data.

Connections

Can we build a physical theory of learning that predicts networks' macroscopic behavior? [Polski]