Back to archive
#ai#llm#glossary#aigen

Positional Encoding

You read “BMW overtook Volvo” and “Volvo overtook BMW”. They contain exactly the same words, but changing their order reverses the cars' roles. If the model sees only word descriptions and compares them as an unordered set, it lacks a signal specifying their places in the sentence.

Positional Encoding supplies information about a text fragment's position. In the original Transformer, a second list calculated from the position number was added to the list of numbers describing a token. A token is a unit of text; its numerical description is called an embedding.

“Volvo” in the first and third positions therefore receives different input signals. The model can learn to use this difference when interpreting the sentence. In the original method, position patterns resembled waves at different frequencies, calculated with sine and cosine functions.

The positional signal does not block access to future fragments. Causal Masking performs that function. Not every model adds position in the same way: Rotary Position Embedding introduces it by rotating parts of the numerical descriptions.

Mechanism source: Explanation: Dive into Deep Learning, §11.6.3.

Mechanism and details

In the original Transformer, a vector built from sines and cosines of different frequencies was added to the representation of each token. The vector had the same dimension as the embedding, so addition proceeded coordinate by coordinate. Attention Is All You Need, §3.5.

What the model receives

The addition can be written as zp=ep+PE(p)z_p=e_p+PE(p): epe_p is the token embedding at position pp, and PE(p)PE(p) depends on position.

In a two-dimensional illustration of the sinusoidal formula, PE(0)=[0;1]PE(0)=[0;1] and PE(1)≈[0,841;0,540]PE(1)\approx[0{,}841;0{,}540]. The same token with the same embedding therefore receives a different input vector at these two positions. These numbers encode position, not the token's meaning or the strength of its relationship with a neighbor.

Positional Encoding does not replace a causal mask. Position information does not forbid access to a future position; the mask does that. The sinusoidal variant is one way to represent position, not a required solution in every model.

See also: Rotary Position Embedding introduces position by rotating query and key instead of adding a position vector to the embedding.