Back to archive
#ai#llm#glossary#aigen

Layer Normalization

One fragment's description enters the next layer with numbers [2, 4], another with [20, 40]. The signal magnitude can differ considerably. You want to aid computation by converting numbers within each description to a comparable scale.

Layer Normalization normalizes a vector — a list of numbers — using its own mean and spread. A typical Transformer does this separately for each text position. It does not need other sentences from the same data group.

For [2, 4], the mean is 3. Subtracting it leaves [-1, 1]; division by the appropriate measure of spread adjusts the scale. Learned coefficients can then stretch and shift each number.

The final result therefore need not have exactly zero mean and unit variance. This is a local operation stabilizing processing, rather than removing differences in meaning between tokens or guaranteeing the stability of the whole network. RMSNorm adjusts scale without subtracting the mean.

Mechanism and details

This distinction is introduced by Ba, Kiros and Hinton, Layer Normalization, §3. Statistics depend on the current input, so normalization does not need other examples in the same batch.

For coordinate xix_i, the expression with learned scaling and shifting is:

yi=γixi−μσ2+ε+βi.y_i=\gamma_i\frac{x_i-\mu}{\sqrt{\sigma^2+\varepsilon}}+\beta_i.

μ\mu is the vector's mean, and σ2\sigma^2 is the mean squared deviation from it. Positive ε\varepsilon protects against division by zero. Parameters γi\gamma_i and βi\beta_i are learned. The Keras documentation gives the full calculation and the meaning of the normalization axis.

An example with two coordinates

For [2; 4], the mean is 3, and the variance is 1. Subtracting the mean gives [−1; 1]. With a very small epsilon, gamma equal to one, and beta equal to zero, the result is close to [−1; 1].

After gamma and beta are learned, the final vector need not have mean zero or variance one. Normalization also does not mean the entire model becomes independent of the scale of all its parameters.

In the original Transformer, LayerNorm followed addition through a Residual connection, for both attention and FFN. This order is shown in Attention Is All You Need, §3.1; it should not be made a rule for all LLMs.