Back to archive
#ai#papers#aigen#llm

Mechanistic interpretability

The model identified Mary as the book's recipient in “John gave Mary a book”. The correct answer does not yet explain how it obtained it. You want to check which internal computations actually caused this result.

Mechanistic interpretability studies mechanisms inside a model that lead to its behavior. Observing that an element is often active when the answer is correct is insufficient. A researcher proposes a hypothesis, changes the element's operation, and observes the effect.

In the example, selected internal values can be replaced with those the model produced for a sentence with the people's roles swapped. These values are called activations — results of intermediate computations. If the answer changes as predicted, the element's role has stronger support than observation alone provides.

The Circuit Condensation study, §2, uses this approach and additionally changes the model to concentrate behavior in a smaller structure. A conclusion always has a scope: particular inputs, the feature being studied, and the intervention type. Explaining one behavior does not fully explain all the model's decisions.