Back to archive
#ai#llm#glossary#aigen

KV Cache

The model appends a response piece by piece. It has already computed much information about earlier text, so computing it from scratch at every step would repeat work. You want to retain results that remain valid.

KV Cache stores keys and values from attention layers. These are lists of numbers used respectively to match fragments and convey their information. It does not store a ready answer or a summary of the conversation.

After processing A, B, and C, a new fragment D uses the stored descriptions of those positions and adds its own. In a model that blocks access to the future, new fragments do not change earlier descriptions, so they can be reused with compatible settings.

We save repeated computation but pay for memory to store the numbers. This cost grows with a long context. Prefill prepares the cache for a known input, and Decode extends it during the response. The required information must still be read and mixed.

Mechanism and details

In a model with Causal Masking, later tokens do not change the representations of earlier positions. This enables reuse of the cache for the same prefix and model settings. The information flow is described in Attention Is All You Need, §3.1 and §3.2.3, and cache updates in each layer in the Hugging Face documentation: How caching works.

What remains from the previous step

An original illustration: the model has processed tokens A, B, and C. When processing D, it uses the stored K and V for the three preceding positions and appends the pair for D. It compares the current position's query with the available keys; the result uses the corresponding values. The cache is neither a ready-made answer nor a textual conversation summary.

Storing results has a cost. A small original example for one layer: four positions, two KV heads, three numbers in each K and V vector, and four bytes per number:

2⋅4⋅2⋅3⋅4=192 bytes2\cdot4\cdot2\cdot3\cdot4=192\ \text{bytes}

The first factor of 2 means separate keys and values. The result includes only tensor elements, without allocation overhead. A full dynamic cache grows with the number of stored positions; limited-window variants can remove older positions.

An older article on KV-cache scaling [Polski] develops the memory-cost topic. The article on moving caches between models [Polski] addresses a separate problem: matching tensor dimensions does not ensure matching meaning. Changing the model or prefix requires reassessing which part of the cache remains valid. See also: PagedAttention describes block-based memory organization, and Prefix Caching describes reusing KV for a shared prefix.