Edge of Stability
Edge of Stability
During learning, a model improves its weights in small steps. You expect its error to decrease smoothly. Yet in some experiments, results oscillate, local conditions appear unstable, and training still progresses. How can this observation be described?
Edge of Stability is a learning regime where the greatest local curvature of the error function stays near the stability boundary for the step being used. Curvature describes how rapidly the guidance on “which direction to improve weights” changes. On a sharp surface, the same step more easily overshoots the minimum.
For ordinary gradient descent — updating weights in a direction that lowers error — the local boundary in a simple quadratic model is 2/η, where η is the learning step size. Cohen and colleagues, §2–3, observed training staying near this boundary and oscillating even while error decreased over a longer period.
This describes dynamics, rather than recommending deliberate destabilization. Results for full gradient descent cannot be transferred to every optimizer and LLM without checking. Scaling Laws concern resource–result relationships; here we ask how learning itself unfolds.
Connections
Can we build a physical theory of learning that predicts networks' macroscopic behavior? [Polski]