PagedAttention
A server handles conversations of different lengths. If it reserves a large contiguous memory region in advance for each entire future response, some space remains unused. You want to allocate smaller pieces as needed.
PagedAttention reads information from KV Cache stored in blocks that may occupy different locations in memory. A block table indicates where to find the next logical part of the conversation. The text remains in the correct order.
If a block holds descriptions of four tokens — pieces of text — a fifth requires a second block. The system does not have to reserve room for each conversation's maximum length in advance. It can also share unchanged blocks according to its memory management rules.
This is a way to address and organize data, rather than automatically shorten context or change the attention result. Tables and partially filled blocks still have a cost. The experiment shows how logical order is separated from physical location.
Mechanism source: Kwon et al., §4.1–4.3.
Mechanism and details

An original example: a block holds KV for four positions. A sequence of length five needs two blocks, or eight slots. The first block is full; the second contains one entry and three free slots. Logical blocks 0 and 1 can point to physical blocks 4 and 1. There is no need to move the existing prefix to append the next block during Decode.
Text order is not a memory address
The demonstration shows one sequence and blocks of four positions. Process successive positions, especially the transition from four to five. Then change the layout of the same data. We assume one scalar head: all keys are zero, the query is 1, and values at successive positions are 1, 2, 3… Attention weights are therefore equal, and the result is the mean of the values. For five positions, it is 3 in both layouts.
This illustrates addressing and allocation, not GPU runtime. Empty slots in the last block still occupy memory; PagedAttention neither removes nor compresses existing KV. FlashAttention focuses on organizing computation and transfers during attention. These mechanisms can complement one another. Block-based memory sharing is also useful in Prefix Caching.
I use AI-generated content as part of my daily learning process.