Back to archive
#ai#llm#glossary#aigen

Cross-entropy

Two models identify the same word as the best continuation, but the first gives it 80% probability and the second only 20%. Checking only the winner misses this difference. Training needs a measure of how strongly the model supports the correct result.

Cross-entropy measures how well predicted probabilities match the target result. With one correct token — a piece of text from the data — the penalty is the negative logarithm of the probability assigned to it. The logarithm makes a very small probability for the correct answer incur a large penalty.

For 0.8, the penalty is about 0.223; for 0.2, about 1.609, using natural logarithms. Training can therefore increase confidence in correct predictions even when the winner is already correct.

Softmax determines probabilities; Cross-entropy evaluates them against a target. A small error on the data does not prove that the model tells the truth in every conversation. The experiment below shows how the assigned probability affects the penalty.

Mechanism and details

H(q,p)=−∑i=1nqiln⁡piH(q,p)=-\sum_{i=1}^{n}q_i\ln p_i

qiq_i is the target probability of candidate ii, and pip_i is its predicted probability. When the label identifies one correct token yy, the entire formula simplifies to −ln⁡py-\ln p_y. Softmax can derive a distribution from logits; Cross-entropy evaluates it against the target. The PyTorch: CrossEntropyLoss documentation describes both types of labels. This particular function accepts logits directly, so they should not be passed through Softmax beforehand.

What the loss value means

An original example: the model assigns the correct token a probability of 0.8. The loss is approximately 0.223. At a probability of 0.2, it rises to approximately 1.609. We use the natural logarithm. Two models can identify the same token as the most probable and still have different losses, because the value assigned to it also matters.

In Autoregressive Language Modeling, the target is the next token in the text. Minimizing the mean of this loss corresponds to maximizing the data log-likelihood, excluding additional regularization terms. This form of the objective is described by Bengio et al., A Neural Probabilistic Language Model, §2.

The average must cover the appropriate positions: padding and tokens excluded from evaluation should not affect it. With soft target labels, the full sum must not be replaced by a single −ln⁡py-\ln p_y. Perplexity is related to the averaged loss for individual tokens.