Multi-Query Attention
The model connects the current piece of text to the earlier conversation through several parallel computations. These computations are called attention heads. If each head stores its own descriptions of the entire preceding conversation, memory and reading costs grow. Can some data be shared while preserving different ways to read them?
Multi-Query Attention (MQA) keeps multiple queries — descriptions of what the current positions need — but gives them a shared key and value set. Keys support matching, while values contain information to convey. All are lists of numbers.
Two heads can use the same key/value set yet assign different weights because they have different queries. Shared data do not force identical results.
The savings concern, among other things, KV Cache, the stored descriptions of earlier text. The change can affect quality, however, and does not imply a proportional speedup of the whole model. Grouped-Query Attention shares data within smaller groups.
Mechanism and details
In ordinary Multi-Head Attention, each head has its own query, key, and value projections. MQA removes the head dimension from K/V and their projections while retaining it in queries. This variant is defined by Noam Shazeer, Fast Transformer Decoding: One Write-Head is All You Need, §3.
The same memory, a different question
An original example has two keys: and . Two query heads receive and respectively. Scaled Dot-Product Attention divides dot products by and applies Softmax. The first head gives the keys weights of about 0.804 and 0.196; the second reverses those proportions. With shared and , these are also the coordinates of two different results.
Where the savings come from
KV Cache stores earlier K/V. With eight query heads, MHA stores eight K/V sets, while MQA stores one. For the same sequence lengths, head dimensions, and precision, the cache tensors themselves are then eight times smaller. This is a size calculation, not an eightfold speedup of the whole model.
Shazeer's analysis in §2.4.1 and §3.1 links this change to memory transfer during successive generation steps. However, the architecture and its representational capacity change; this is not lossless removal of identical copies from an arbitrary trained MHA. Grouped-Query Attention retains several K/V sets as an intermediate variant.
I use AI-generated content as part of my daily learning process.