Back to archive
#ai#llm#training#papers#aigen

RMSNorm

One layer receives a list of large numbers, another much smaller values. These scale differences can complicate computation and learning. You want to standardize the signal's overall magnitude while preserving proportions between its elements.

RMSNorm divides a list of numbers by its RMS, the square root of the mean of their squares. It can then multiply each coordinate by a separate learned coefficient. The list is a vector; in a model, it describes the piece of text being processed.

For [3, 4], RMS is about 3.54. Dividing gives about [0.85, 1.13]: the shared scale decreased while the proportion stayed the same. Implementations add a small constant to protect against division by zero. Zhang and Sennrich, §3, describe this mechanism.

Unlike Layer Normalization, RMSNorm does not subtract the input numbers' mean. Normalizing one branch alone does not remove scale from all model computations, particularly the Residual connection path that forwards the input. This is a local operation, rather than a guarantee that the entire network is invariant.

Connections

Learning rate is not enough. ELR matters [Polski]

Scale invariance