Back to archive

Scaling Laws

Scaling Laws

You have a budget for model training. You can build a larger model or show a smaller one more text. Fully testing every combination would be expensive. You need a way to estimate what extra scale provides before spending the entire budget.

Scaling Laws are empirical relationships between model size, data quantity, computation, and training results. “Empirical” means discovered through measurements. They often describe a decline in text prediction error as resources increase.

Researchers train several smaller models and fit a curve. They use it to try to predict the result of a larger training run. Doubling the number of parameters, the model's learned numbers, need not double quality: the benefit also depends on data quantity. Kaplan and colleagues, §3–6, describe these relationships; the Chinchilla paper studies budget allocation between model and data.

This is a planning tool, rather than a law guaranteeing a larger model performs every task better. Lower text prediction loss does not automatically mean more truthful answers. Changing the data, architecture, or training objective may require new measurements.

Connections

Can we build a physical theory of learning that predicts networks' macroscopic behavior? [Polski]