Back to archive
#ai#llm#glossary#aigen

Prefill

You send a long question to a chat and wait for the first word. Before the model appends a response, it must process the content it has already received. This initial stage can require considerable work, especially for a long document.

Prefill processes the known input, called the prompt, before further generation. The model creates numerical descriptions of tokens — pieces of text — and can retain some results in KV Cache for reuse.

Because the entire input already exists, many positions can be computed in parallel within a layer. Context access rules still apply: parallel work does not let an earlier position look at a prohibited future.

The final position's result allows the probabilities of the first new token to be computed. Choosing it is a separate step. Prefill does not cover all the time since sending the question: network time, queuing, and other stages also contribute. Decode describes further extension of the response.

Mechanism source: NVIDIA, Mastering LLM Techniques: Inference Optimization, section “Understanding LLM inference”.

Mechanism and details

Several prepared paper strips pass together under a wide comb, with the last reaching an empty space for a new element.

A known prompt can be processed in parallel

All input tokens are already available, so many positions can be calculated simultaneously within a layer. Causal Masking still applies: the second position does not use the third. Parallel computation does not remove this dependency or the order of traversal through layers. The foundation is described by Vaswani et al., §3.1–3.2.

Prefill can also process the prompt in chunks [Polski], retaining the cache of earlier positions. vLLM documentation, “Chunked Prefill” describes using this division to interleave prompt processing with response generation for other requests.

Four positions, the same result

An original example isolates one attention head with dimension 1: all queries and keys are 1, and the values are [2,4,8,10][2,4,8,10]. Allowed comparison scores are equal, so each position returns the mean of its visible values: [2,3,14/3,6][2,3,14/3,6]. Compare processing all four together with chunks of two. In the second variant, the third position must still see the first two.

This calculates a single operation, without trained model weights or GPU measurements. In a full model, the cache is maintained separately for each layer. Changing the chunking may introduce small numerical differences in the implementation; it does not change the mathematical mask.

After the prompt comes Decode, where successive inputs depend on newly selected tokens. Prefill time alone does not cover the user's entire wait: the queue, input preparation, and delivery of the result also matter. See also: Time to First Token covers the client's wait for the first token, and Continuous Batching allows the group being served to change between iterations.

I use AI-generated content as part of my daily learning process.