Back to archive

Neural Feature Ansatz

Neural Feature Ansatz

A model recognizes an image, but we do not know which input changes actually affect the answer. Do learned weights amplify precisely the directions of change to which the result is most sensitive? This asks how the model selects features.

Neural Feature Ansatz (NFA) is a hypothesis connecting weight structure with output sensitivity to input. “Ansatz” means a proposed relationship that needs justification or testing. Weights are numbers controlling computation, and their arrangement can be compared with directions important for prediction.

An intuitive example: if a small change in an image fragment strongly changes the result, that direction may be an important feature. Researchers collect these local sensitivities, called gradients. They form a table of numbers describing which directions of input change jointly affect the result. They compare it with a table calculated from network weights. The hypothesis predicts a specific relationship between these tables; it is more than asking “is the model looking at the right place?”.

The feature-learning paper, §3, presents NFA and how to test it. It is neither a general law for every network nor a claim that every layer behaves identically. The hypothesis's scope depends on the system studied.

Connections

Can we build a physical theory of learning that predicts networks' macroscopic behavior? [Polski]