Softmax
The model scored three possible sentence continuations with the numbers 2, 1, and 0. You want to sample a response according to their probabilities. Raw scores are not probabilities: they can be negative and do not sum to one.
Softmax converts these scores, called logits, into positive shares summing to one. It first applies the exponential function — a transformation that grows with the score — and then divides each value by their combined sum.
For scores 2, 1, 0, we obtain about 67%, 24%, 9%. The first candidate has the greatest probability, but sampling can still select the second or third. Softmax itself neither samples nor appends text.
The same computation can determine fragments' shares in attention. Those shares concern mixing information, rather than choosing the next token. Even a token probability of 67% is not a judgment of the truth of the entire statement.
Mechanism and details
is the logit of candidate , is the number of candidates, and is its probability. In a tensor, the normalization axis must be specified: the sum is one for each group considered along that axis. The formula and the meaning of the dim parameter are given in the PyTorch: Softmax documentation.
A small example
Suppose there are three logits: . Exponentiation gives , and dividing by the sum gives . This is an original numerical example. The first candidate has twice the probability of each of the others.
In a language model, LM head produces logits for the vocabulary. Softmax converts them into a distribution, but does not select a token. In the example, selecting the maximum always chooses the first candidate; sampling from the distribution can select any of the three.
In Scaled Dot-Product Attention, the same operation normalizes query–key comparison scores. The components are then position weights rather than probabilities of vocabulary tokens. Both uses are described in Attention Is All You Need, §3.2.1 and §3.4.
A value of 0.5 in the token distribution does not mean a 50% chance that the whole sentence is true. The result concerns a choice among specific candidates for a given input, not an assessment of a statement's truth.