Back to archive
#ai#llm#glossary#aigen

Self-attention

In “Anna put the book down because it was heavy”, the word “it” alone explains little. It must connect to earlier fragments. A model processing text needs a way for one position's description to incorporate information from other positions in the same text.

Self-attention creates such a description by mixing information from available tokens, pieces of text, with weights. For each position, the model prepares three lists of numbers: a query to find relationships, a key for comparison, and a value containing information to convey. They are learned representations, rather than literal questions or dictionary entries.

For the position “it”, its query is compared with keys of other fragments, such as “Anna” and “book”. The results determine the shares of information conveyed from their values. If the book fragment receives a larger share, its description has more influence on the new description of “it”. “Self” means all three kinds of data come from the same sequence.

The sentence example shows the need for context, but does not guarantee that a particular computation correctly identifies the pronoun's reference. Causal Masking can further restrict available positions. Attention weights do not assess the truthfulness of the entire response.

Mechanism and details

A position's query is compared with the keys of available positions. The resulting weights determine the contribution of their values to the result. In a Transformer, this calculation is performed by, among other methods, Scaled Dot-Product Attention. Description and implementation: Dive into Deep Learning, §11.6.1.

What limits access to other positions?

Take three positions in a sequence. Without a mask, the second position's representation can use all three, including itself. With a causal mask, the third position is unavailable to it. This is still Self-attention: the scope of access changed, not the source of query, key, and value.

This is how the encoder and decoder were distinguished in the original paper, §3.2.3. The word “self” alone therefore does not imply access to future tokens.

In the full variant, the number of compared pairs grows quadratically with sequence length. There are 10 000 pairs for 100 positions, and 40 000 for 200. This illustrates pair counts, not the runtime of a specific implementation. Cost comparison: §11.6.2.