Back to archive
#ai#llm#glossary#aigen

Top-k Sampling

When sampling the next piece, the model allows many extremely unlikely options. You want to keep some diversity but remove the long tail of candidates. You can admit only a specified number of leaders.

Top-k Sampling samples the next token, a piece of text, from the k highest-scoring candidates. It excludes the rest and renormalizes the retained probabilities to sum to one.

If A has probability 0.5, B 0.3, and the rest together 0.2, k=2 keeps A and B. After dividing by 0.8, their probabilities are 0.625 and 0.375. B can still win; the method does not give everyone equal chances.

A fixed k determines the number of candidates, rather than their total credibility. With a very flat distribution, even the leaders may have small probabilities. Top-p Sampling determines the set by cumulative probability, while Temperature adjusts the proportions.

Mechanism and details

A practical implementation can filter logits before Softmax: discarded positions are assigned negative infinity. Their probability after normalization is then zero. This mechanism is described in the Transformers: TopKLogitsWarper documentation.

Two candidates does not mean equal chances

An original example for k=2k=2:

TokenBefore filteringAfter filtering and normalization
A0.500.625
B0.300.375
C0.150
D0.050

A and B had a combined probability of 0.80. We divide by this sum: 0,50/0,80=0,6250{,}50/0{,}80=0{,}625 and 0,30/0,80=0,3750{,}30/0{,}80=0{,}375. Sampling can still select B. With k=1k=1 and a unique maximum, we always select the highest-scoring token.

A fixed kk controls the number of candidates, not their combined probability before the cutoff. Two retained positions can contain almost the entire distribution or only a small part of it. Top-p Sampling sets the boundary according to the sum of probabilities. Temperature changes the relative chances; positive scaling alone does not change the order of finite logits.

For ties, it is worth checking the implementation. The Transformers 5.17.0 code removes scores strictly below the score of the k-th candidate. It can therefore retain additional tokens with the same value at the boundary. The example in the table has no ties.