LM head
The model has processed “Today I drink” and has an internal description of this content as a list of numbers. The user, however, is waiting for text. We need a transition from the description hidden in the model to scores for candidates for the next response fragment.
The LM head is a language model's final layer, computing a score for every token in its vocabulary. A token is a piece of text, and its score before conversion into probability is called a logit.
In a simplified example, the input description is (2, 1). A candidate with weights (1, 0) gets 2, while one with (0, 1) gets 1: we multiply corresponding numbers and add the results. Softmax converts scores into probabilities. Only the generation rule then chooses a token — the highest-scoring one or a sampled one.
The LM head itself does not check whether an answer is true. A small change to the hidden description can also change the winner if candidates are close. The original Transformer, §3.4, describes projecting the output onto the vocabulary. The article on quantization shows how differences in earlier computations reach this stage.