Top-p Sampling
Sometimes the model has one clear leader; sometimes it has many continuations with similar scores. Always admitting five candidates ignores this difference. You want to choose the group's size according to where the probability is concentrated.
Top-p Sampling, or Nucleus Sampling, chooses the smallest leading group of candidates ordered by descending probability whose sum reaches threshold p. It then samples from that group after renormalizing the probabilities to one.
For probabilities 0.7; 0.15; 0.1; 0.05 and threshold 0.75, the first two remain: together they have 0.85. For 0.4; 0.3; 0.2; 0.1, three are needed. Each generation step can therefore admit a different number of tokens — pieces of text.
The threshold does not mean a minimum probability for each candidate or guarantee answer quality. Top-k Sampling fixes the number, while top-p responds to the distribution; neither filter assesses the truthfulness of the text itself.
Mechanism and details
The method and normalization are described by Holtzman et al., The Curious Case of Neural Text Degeneration, §3.1. The candidate set is recalculated at every step of Autoregressive Language Modeling.
The same threshold, a different number of tokens
An original example for ; the rows already contain sorted probabilities:
| Distribution for A, B, C, D | Retained tokens | Their sum before normalization |
|---|---|---|
| 0.70; 0.15; 0.10; 0.05 | A, B | 0.85 |
| 0.40; 0.30; 0.20; 0.10 | A, B, C | 0.90 |
In the second row, A and B total only 0.70, so C must be included. After normalization, their chances are , , and respectively. D is excluded. The threshold does not mean retaining tokens each of which individually has a probability of at least 0.75.
Top-k Sampling would try to retain the same number of candidates in both cases. Top-p responds to the distribution of their chances. If we apply Temperature beforehand, the sum of successive probabilities changes, and the cutoff may fall in a different place.
The paper studied, among other things, open-ended text continuation by GPT-2 (§4.1). Its results do not establish one optimal threshold for every model and task. Nor is the parameter the probability that the generated response will be correct.