Cross-attention
In translation, the model has two different texts: the original and the response still being created. When appending another translation word, it needs information from the original. Mixing only the response already written cannot provide missing source content.
Cross-attention lets positions in one sequence read information from another. A query — a list of numbers specifying the current position's need — comes from the response being created. Keys match source fragments, and values contain their information to convey.
Each response position receives its own mixture of information from the original. Two response positions and three source fragments give two results, one for each reading position, rather than three.
In Self-attention, all these descriptions come from the same sequence. “Cross” indicates different sources, rather than necessarily two independent models. The weights need not be a literal, correct word-to-word alignment between languages.
Mechanism and details
In the original Transformer, queries came from the decoder side, while keys and values came from the encoder's output. The paper calls this part “encoder-decoder attention.” Attention Is All You Need, §3.2.3.
A result is produced for each query
Suppose the decoder supplies two query vectors and the encoder three key–value pairs. Each query then has three possible connections to the source. Two result vectors are produced, one for each query position, rather than three vectors corresponding to the encoder's input length.
In translation, this can be understood as follows: the representation of the response being created determines how to combine information from the source sentence's representation. This illustrates information flow, not a guarantee that an individual attention weight is a correct word alignment between languages.
The CrossAttention implementation in TensorFlow passes the query sequence and key–value context separately. The operation itself does not transfer information directly between query positions; each reads the context.
Cross-attention can use Multi-Head Attention. “Cross” refers to the sources of representations, while “multi-head” refers to multiple sets of projections. The two sequences need not come from two independently trained models: encoder and decoder may be parts of one model.