Grokking
Grokking
The model answers training tasks perfectly but makes mistakes on new examples of the same kind. It looks as though it memorized a table of answers. Training continues, and only much later do results on new examples improve sharply.
This delayed improvement in generalization, performance on unseen data, is called Grokking. The distinction between quick success on the training set and later success outside it matters.
In a simple arithmetic task, the model may first fit answers for known pairs of numbers and only later learn a structure that handles new pairs. Power and colleagues, §2–3, study this pattern in small algorithmic tasks. Discovering a rule is an intuition for interpreting the result, rather than proof of human “understanding” in the model.
Recognizing grokking therefore requires two measurements: performance on training data and on separate evaluation data. A decline in training error alone is insufficient. The phenomenon does not occur in every training run or promise that any overfit model will improve if we simply wait longer.