Decode
The first piece of the response has appeared. The next must account for what was just written: after “Today I drink tea” the model has a different context than after “Today I drink water”. This changing text must be extended step by step.
Decode is the stage of generating a continuation after processing the initial input. In the usual loop, the model processes the most recently selected token — a piece of text — computes probabilities for the next one, and a selection rule appends one of the candidates.
The chosen piece becomes the input to the next step. KV Cache retains needed results from earlier computations, so the entire history does not have to be processed again.
Successive choices depend on each other, although several different responses can be served together. Here Decode means extending text with the model, rather than merely converting token numbers into characters. Prefill prepares the initial context.
Mechanism source: Shazeer, Fast Transformer Decoding, §2.4.
Mechanism and details

The first response token can be selected from the output of Prefill. If generation continues, this token becomes the input of the next model pass. Its new K and V are added to KV Cache, and the query uses the current and earlier positions. Selecting the token ID alone does not yet append its representation to the cache. This can be seen in the Hugging Face generation loop, “Cache storage implementation”.
A token selected but not yet processed
An original miniature model has three tokens: A, B, C. Their vectors are , , and respectively. We take equal to these vectors; attention uses a scale of . From the result , we create logits , and Greedy Decoding selects the largest.
The prompt A B gives approximately , so the first answer is A. The cache still contains only the prompt's two positions. Only the next pass processes the selected A.
This is a fully specified demonstration model, without training, positions, multiple layers, or an end token. Its repetitive response has no linguistic meaning. Comparison with recomputing the entire prefix checks agreement of results, and the counter covers only positions for which K/V were calculated.
The cache removes repeated computation of earlier K/V, but attention still reads the context. With small batches, transferring weights or cache often becomes the bottleneck; this depends on the model, hardware, and sequence length. In Disaggregated Serving [Polski], Decode runs on separate resources after receiving the cache. Decode in this sense is not a tokenizer converting token identifiers back into text.
I use AI-generated content as part of my daily learning process.