Prefix Caching
You ask five questions about the same long document. Each input starts with the same instruction and document text; only the ending contains a new question. Computing the same beginning again would repeat work.
Prefix Caching reuses previously computed results for a common prefix of tokens — pieces of text — across requests. In suitable models these may be KV Cache data that let the model skip part of input processing.
The new question still needs processing, and its answer still needs generating. The cache is not a library of ready answers. The prefix and computational conditions must match: two documents with similar meanings need not share a cache.
Even a small change to an earlier part can prevent reuse of later blocks. The saving concerns repeated work, rather than every conversation; storing shared results also takes memory.
Mechanism source: vLLM 0.21.0, “Introduction” and “Limits”.
Mechanism and details

What matters is an identical token prefix in a compatible computational context, not similarity of meaning. Keys and values in later layers depend on the preceding context. The same final paragraph after a changed beginning does not automatically yield the same KV.
vLLM identifies a block by its tokens, the preceding block's identifier, and additional data, including the LoRA adapter ID and image input. In the version described, it stores full blocks for reuse. See the description of Automatic Prefix Caching. Compatibility of the model, positions, and other settings affecting KV is a condition for correct reuse; comparing texts alone is insufficient.
One changed token, different consequences
The cache in our example contains the sequence A B C D E F G H, divided into blocks of two tokens. Letters represent illustrative token IDs. A new request replaces exactly one position with X. Change its location and check how many whole blocks can be recovered.
Changing H to X leaves seven shared tokens, but only six in full blocks. Changing A to X gives zero hits, even though the next seven tokens match. All other conditions in the example stay fixed; the cache is available, and nothing has been evicted.
This caches computation, not a ready-made answer. Decode must still produce further tokens, and evicted blocks must be recalculated. PagedAttention explains organizing memory in blocks.
I use AI-generated content as part of my daily learning process.