Critical Batch Size
Critical Batch Size
You want to speed learning by showing the model more examples at once. Initially this helps. Beyond a point, however, you process much more data while the number of steps needed for a given result decreases only slightly. Where does enlarging the group stop paying off?
Critical Batch Size is the characteristic batch scale beyond which further increases in the data group per step yield diminishing benefits. A batch is the examples or tokens — pieces of text — used together to compute one model weight update.
Averaging more examples reduces randomness in the guidance on improving weights. But with a fixed amount of data, a larger batch means fewer updates. Moving from 100 to 200 examples may save many steps; another doubling need not give a proportional gain. This illustrates the tradeoff, rather than giving universal boundary values.
McCandlish and colleagues, §2–3, model this relationship using the gradient noise scale. The boundary depends on the task and training phase. Faster data processing on hardware need not mean equally effective learning per token read.
Connections
Can we build a physical theory of learning that predicts networks' macroscopic behavior? [Polski]