µP and Hyperparameter Transfer
µP and Hyperparameter Transfer
Choosing a good learning step requires trials. For an enormous model, each trial is expensive. Can good settings be found in a small network and used after increasing its width, the number of elements in its layers?
Maximal Update Parametrization (µP) scales parameters and their updates as network width changes. It aims to retain comparable learning behavior, allowing some hyperparameters — settings chosen before training — to transfer from smaller models.
In practice, we first prepare a small model according to µP, test several learning step sizes, and then enlarge the model while preserving the appropriate scaling rules. This is more than copying one number into an arbitrary larger network: initialization and parametrization of individual layers also matter.
Yang and colleagues, §2–3, describe settings transfer and tests of this method. The saving is cheaper tuning, but transfer has a specified scope. Changes to data, architecture, depth, or learning rules may require separate checks. See also the later paper cited in the original note.
Connections
Can we build a physical theory of learning that predicts networks' macroscopic behavior? [Polski]