Back to archive
#ai#llm#glossary#aigen

Scaled Dot-Product Attention

The model needs to combine information from several sentence fragments. One is more useful, another less so. It needs a specific way to compute their contributions and then assemble the information into one result.

Scaled Dot-Product Attention compares numerical query and key descriptions, converts results into weights, and mixes the corresponding values. Each description is a vector, a list of numbers. A query describes a position's current need, a key supports matching, and a value contains information that can be conveyed.

Comparison multiplies corresponding numbers and sums the products. The result is divided by the square root of the list's length so its scale does not grow excessively with dimension. Softmax converts results into nonnegative shares summing to one.

If the shares are 25% and 75%, and the numbers being conveyed are 2 and 6, we obtain 0.25 × 2 + 0.75 × 6 = 5. This simplifies the operation to one number; real values are lists. The shares describe information mixing, rather than the probability that an answer is correct.

Mechanism and details

Without masking and dropout, the expression is:

Attention⁡(Q,K,V)=softmax⁡(QKTdk)V\operatorname{Attention}(Q,K,V)=\operatorname{softmax}\left(\frac{QK^T}{\sqrt{d_k}}\right)V

QQ contains queries, KK contains keys, and VV contains values. Vectors occupy matrix rows; dkd_k is the dimension of query and key. Softmax operates separately for each query over available keys. The matrix of its results is used to combine values. Derivation and implementation: Dive into Deep Learning, §11.3.3.

A small example of the result

Suppose softmax produces weights 0,250{,}25 and 0,750{,}75, and the two values are [2;0][2;0] and [0;4][0;4]. The result is:

0,25[2;0]+0,75[0;4]=[0,5;3].0{,}25[2;0]+0{,}75[0;4]=[0{,}5;3].

This is an invented numerical example. The weights distribute the contribution of values; they are not probabilities that the model's answer is correct.

Dividing by dk\sqrt{d_k} counteracts the growth in dot-product scale with dimension. Excessively large values can saturate softmax and reduce its gradients. The rationale is given in Attention Is All You Need, §3.2.1.

This operation does not determine where the data comes from. In Self-attention, query, key, and value come from one sequence; in attention between sequences, their sources differ. Multi-Head Attention performs these calculations on several learned projections.