Back to archive
#ai#llm#training#papers#aigen

Weight decay

The model fits the data increasingly well, while its weights — numbers controlling computation — move away from zero. You want to constrain this growth without manually correcting millions of values. You need a rule operating at every training step.

Weight decay pulls weights toward zero. In the variant separated from the learning algorithm's main update, each weight is multiplied by a coefficient slightly below 1 and then changed according to information from the data.

With a shrinking factor of 0.99, weight 2 becomes 1.98 and weight −2 becomes −1.98 before accounting for the learning update. The distance from zero therefore shrinks, rather than always the numerical value itself. This does not mean every weight ultimately approaches zero: the data update can act in the opposite direction. The amount of shrinking depends on the learning step and decay coefficient. Loshchilov and Hutter, §2, explain separating weight decay from the gradient-based change.

Controlling weight scale can affect both fitting new data and training dynamics. The effective learning step study, §4, analyzes the latter role. This is a result under the studied conditions, rather than proof that smaller weights always make a better model.

Connections

Learning rate is not enough. ELR matters [Polski]

See also: AdamW [Polski] — applying weight decay outside Adam [Polski]'s step-selection mechanism.