Back to archive
#ai#llm#glossary#aigen

Multi-Head Attention

One combination of information can mix different sentence relationships too strongly. In translation, relationships involving people, time, and grammatical form may all be useful. We want to let the model compute several different ways of combining the same context.

Multi-Head Attention (MHA) performs several attention computations, called heads. Each uses its own learned transformations of queries, keys, and values: lists of numbers used to match and convey information. All can operate on the same tokens, pieces of text.

In an illustrative model, three heads each return four numbers. Their results are placed side by side in a list of twelve numbers and then transformed together. This is neither a simple vote nor an average of three answers.

The design allows different mixing patterns, but does not assign roles such as “grammar head” or “time head” in advance. What a head actually learned must be checked. Self-attention specifies the data source, while MHA specifies the number of parallel computational sets; both concepts can occur together.

Mechanism and details

In a Transformer, individual heads perform Scaled Dot-Product Attention. Their results are then combined by concatenation — placing vectors side by side — and transformed by a shared output projection. This is not simply averaging the results. Construction and implementation: Dive into Deep Learning, §11.5.

Several sets of weights for the same sequence

In an illustrative setup, the representation has 12 coordinates, and three heads each return four. Concatenation gives 12 coordinates, which the output projection can mix further. Each head calculates its own weights for available positions. The sentence need not be divided into three disjoint parts: the projections of representations differ, not necessarily the token sets.

The mechanism enables different relationships to be used in parallel but does not assign roles such as “this head handles grammar.” Its specific behavior comes from training and must be examined.

MHA and Self-attention describe different features of the calculation. The former refers to multiple heads, the latter to query, key, and value coming from the same sequence. They can occur together. In the original architecture, MHA was also used between encoder and decoder. Vaswani et al., §3.2.2–3.2.3.

See also: Multi-Query Attention shares one K/V set among query heads, while Grouped-Query Attention shares K/V within groups.