Back to archive
#ai#llm#training#papers#aigen

Delayed acceleration

You compare two training runs. In the first, you constrain weight growth; in the second, you do not. Initially the first performs worse, so it looks like a poor choice. Near the end, the situation reverses: it catches up and overtakes the second.

Delayed acceleration names a benefit that becomes visible only in a later training phase in a study of controlling weight magnitude — the numbers controlling a model's computations. Performance is assessed through loss, a number measuring model error; a lower value means a better fit to the task being studied.

The authors of the ELR study, §4, relate this pattern to the effective learning step. As weights grow, a change of a given magnitude can become relatively smaller. Weight decay or Hyperball controls that scale. Larger relative changes can both speed acquisition of useful signal and increase fluctuations that initially obscure the benefit.

This explains a specific observed pattern, rather than promising that every worse curve will eventually become better. Comparing methods requires a sufficiently long training run and attention to its settings. Ending measurement early can change the assessment of the experiment.

Connections

Learning rate is not enough. ELR matters [Polski]

Weight decay

Hyperball