Back to archive

Implicit Regularization Through Hyperparameters

Implicit Regularization Through Hyperparameters

Two models train on the same examples and reach similar training error, but one handles new data better. Only their training procedures differed. Where can the difference come from if the objective was the same?

Implicit Regularization is the effect of the learning procedure itself on which solution is selected. Many weight settings can fit the data. Learning step size, examples per step, and randomness in their selection can make training reach some solutions more often than others.

Imagine several routes leading to equally good practice results. Changing how you travel along them may lead to a different place that behaves differently on a new test. This is an analogy for selecting a solution, rather than adding extra knowledge to the model.

Smith and colleagues, §2–3, analyze this effect in stochastic gradient descent [Polski], learning with randomly selected example groups. Not every settings difference improves generalization. Weight decay explicitly controls weights; in simple variants it corresponds to an extra penalty, so not all regularization should be called “implicit”.

Connections

Can we build a physical theory of learning that predicts networks' macroscopic behavior? [Polski]