Back to archive
#ai#llm#training#papers#aigen

Frobenius norm

You have a weight table with thousands of numbers and want one number describing their overall magnitude. A simple sum is misleading: 3 and −3 cancel even though both are large. You need a measure accounting for all entries without this cancellation.

The Frobenius norm is the square root of the sum of squares of all numbers in a matrix. A matrix is simply a rectangular table of numbers. Squaring makes positive and negative entries both increase the result.

For a table containing 3, 4, and two zeros, we obtain √(3² + 4²) = 5. This is the same calculation as a vector's length, applied to all table cells. The NumPy documentation gives this definition.

Multiplying every weight by two also doubles the norm. In Hyperball, it keeps weight magnitude fixed; in the study of the effective learning step, it is part of the update-scale indicator. The norm alone tells us neither how good the model is nor which particular weights are useful: two different tables can have identical norms.

Connections

Learning rate is not enough. ELR matters [Polski]