Back to archive
#ai#llm#glossary#aigen

Perplexity

You compare two models on the same text. You want to know which predicts its successive fragments better, rather than merely which guessed more individual winners. You need a number summarizing the probabilities assigned to the actual sequence.

Perplexity (PPL) summarizes how well the model predicts successive tokens, pieces of the evaluated text. The smaller the probabilities it assigns to fragments that actually occurred, the higher the score. A lower score means better prediction of that text. The formula below, after the example, shows the exact calculation.

If each correct token has probability 1/4, PPL is 4; with probability 1/2, it is 2. This gives an intuition for effective uncertainty, but is not a literal count of words considered at each step.

Comparison requires compatible data, tokenization, and context access. PPL does not directly measure an answer's truthfulness or usefulness. The “Penguin in the fridge” experiment shows how one highly unexpected token affects the score for the entire text.

Mechanism and details

PPL=exp⁡(−1N∑t=1Nln⁡p(xt∣x<t))\mathrm{PPL}=\exp\left(-\frac{1}{N}\sum_{t=1}^{N}\ln p(x_t\mid x_{<t})\right)

NN is the number of evaluated tokens, and x<tx_{<t} is their preceding context. With the natural logarithm, this is exp⁡(H)\exp(H), where HH is the mean Cross-entropy for individual correct tokens, without additional loss terms. Equivalently, it can be calculated as the geometric mean of the reciprocals of those probabilities. This definition is given by Bengio et al., §2.

How to read the result

An original example: a model assigns every evaluated token a probability of 0.25. PPL is then 4. When each of those probabilities rises to 0.5, PPL falls to 2. This does not mean the model always chooses from exactly four or two tokens; the result summarizes the entire evaluated text.

A penguin in the refrigerator

You open the refrigerator. A penguin is sitting inside. For you, this is a logistical problem; for the language model, it is the final token in a sentence. Are nine very good predictions enough to cover the cost of one such surprise?

This is an original thought experiment: two invented models, manually assigned probabilities, and ten illustrative tokens. Each model evaluates the same text, token by token, knowing the preceding tokens of that text. We show only the probability of the token that actually occurred; the remaining mass belongs to other possibilities. These are not measurements of a particular LLM.

The interactive experiment requires JavaScript. An example calculating the result is provided below.

In the initial setting, model A assigns 90% to each of the first nine tokens and 0.01% to the last one. Model B assigns 50% to every token. The ordinary mean probability is about 81% for A and 50% for B, but B wins on PPL: 2 versus approximately 2.762 for A.

Why the reversal? Expressing the cost as −log⁡2p-\log_2 p, the first nine tokens cost model A a total of about 1.368 bits. “pingwina.” alone costs about 13.288 bits. Dividing the sum by ten gives 1.466 bits per token, or PPL≈2,762\mathrm{PPL}\approx 2{,}762. Model B pays one bit per token. A large surprise sharply increases the average logarithmic loss, even when the other tokens had high probabilities. Try raising the chance of “pingwina” to 1%: A wins again.

What can be compared?

A comparison requires the same data and consistent calculation rules, including tokenization, such as Byte Pair Encoding. The available context also matters. Splitting text into disjoint segments deprives the model of preceding tokens at the start of each segment and may worsen the result. The Hugging Face documentation describes evaluation with a sliding window and explains why ordinary PPL is not applied directly to BERT.

Lower PPL means better prediction of the evaluated text under the given procedure, but does not guarantee a better response to a particular task. In Attention Is All You Need, §5.4, label smoothing worsened PPL while improving accuracy and BLEU.