Temperature
The model usually chooses similar continuations, and you want to regulate sampling diversity. You do not change its knowledge or weights. You change only how strongly the best-scoring candidates' lead affects their probabilities.
Temperature is a parameter that divides candidate scores, called logits, before Softmax converts them. A smaller positive value concentrates probability on the leaders; a larger one makes probabilities more even.
In an illustrative example with two candidates, probabilities of 80% and 20% at T=1 become about 67% and 33% at T=2. The second can be sampled more often, although it remains second. A positive temperature does not change the ordering of fixed scores.
Therefore, selecting only the maximum, without sampling, does not change the winner. Temperature is not a measure of creativity or truthfulness either: it increases diversity among attempts, which can bring both interesting and unsuitable responses. The demonstration shows only the change in proportions.
Mechanism and details
is the logit of candidate , and is the number of candidates. For , we get ordinary Softmax. A smaller positive strengthens the advantage of the highest logits, while a larger value flattens the distribution. This formula is presented by Hinton, Vinyals and Dean, Distilling the Knowledge in a Neural Network, §2, in the context of learning through distillation.
An example with two candidates
Take the logits . At , the probabilities are . Increasing to 2 gives . The second candidate has a greater chance of being sampled, though it remains in second place. This is an original numerical example.
In Autoregressive Language Modeling, this scaling can be applied at each generation step, after LM head produces the logits. The Transformers: TemperatureLogitsWarper documentation describes this mechanism for sampling and notes that it requires do_sample=True in generate.
For fixed logits, dividing by a positive preserves their order. Simply selecting the maximum, without sampling, will therefore identify the same candidate. The formula is undefined for ; zero should not be substituted into the denominator.
Changing Temperature during generation does not learn new weights or check whether statements are true. It controls the probabilities of the next choice, so greater response diversity alone does not demonstrate their quality.