Grouped-Query Attention
The model reads earlier text in several ways in parallel. Each such computation is an attention head. You want to reduce the memory these heads occupy, but sharing a single data set completely may constrain the model too much. An intermediate variant is possible: some heads use one shared set, and others use another.
Grouped-Query Attention (GQA) divides query heads into groups sharing key and value. A query is a numerical description used to read context, a key is used for matching, and a value conveys information. Each head still computes its own weights.
Eight query heads divided into two groups of four mean two key/value sets, rather than eight. One group gives MQA; one group per head corresponds to ordinary MHA.
Fewer sets reduce the cost of storing KV when dimensions match. This is a tradeoff in design and quality, rather than a switch guaranteeing a particular speed. Time also depends on hardware and the remaining computations.
Mechanism and details
For query heads and groups, we get K/V sets. When , this is Multi-Query Attention. When , we return to Multi-Head Attention. This family and the GQA-g notation are described by Ainslie et al., GQA, §2.2.
Calculate memory before guessing speed
An original calculation for KV Cache assumes batch size 1, 32 layers, 1024 stored tokens, a dimension of 128 for each key/value head, and 2 bytes per number. The K/V elements alone occupy:
The first factor of two denotes K and V. For eight query heads, the result is 128 MiB in MHA, 32 MiB for two KV heads, and 16 MiB in MQA. A MiB is bytes. We do not count model weights, buffers, memory alignment, or replication between devices.
Reducing the number of KV heads here does not shorten the context or remove query heads. It limits the number of separate K/V representations. The memory calculation does not predict quality or execution time; these also depend on training, hardware, and implementation.
When converting a checkpoint, the authors averaged K/V projections within groups and then continued Pre-training (§2.1–2.2). The results in §3 apply to specific T5 variants and tasks. They establish neither a universal number of groups nor a guarantee of preserving quality by simply switching the configuration.
I use AI-generated content as part of my daily learning process.