Back to archive
#ai#llm#glossary#aigen

Rotary Position Embedding

The model compares pieces of text, but also needs to know how far apart they are. “Yesterday” beside a verb can play a different role from the same word in a distant quotation. How can position be included when comparing numerical descriptions?

Rotary Position Embedding (RoPE) rotates pairs of numbers in queries and keys according to the position of a token, a piece of text. Queries and keys are vectors — lists of numbers compared in attention.

Imagine two arrows with specified lengths. Position determines their rotation angles. The comparison depends on the difference between angles, and therefore the difference between positions. A real model rotates many pairs at different frequencies.

This is an analogy for operations on numbers, rather than text moving in memory. RoPE does not prohibit access to the future or guarantee good performance beyond lengths seen in training. It differs from adding a positional signal in the original Positional Encoding.

Mechanism and details

Unlike vector addition in the original Positional Encoding, this changes the orientation of query and key. For a two-dimensional pair u=(a,b)u=(a,b), rotation by angle ϕ\phi gives:

R(ϕ)u=(acos⁡ϕ−bsin⁡ϕ, asin⁡ϕ+bcos⁡ϕ)R(\phi)u=(a\cos\phi-b\sin\phi,\ a\sin\phi+b\cos\phi)

Rotation preserves vector length. Su et al., RoFormer, §3.1–3.2 derive the construction and combine many such pairs with different frequencies.

Shift both tokens and check the result

If the unrotated query and key stay fixed, then:

(R(mθ)q)⊤R(nθ)k=q⊤R((n−m)θ)k(R(m\theta)q)^\top R(n\theta)k=q^\top R((n-m)\theta)k

mm and nn are positions, and θ\theta is the angle per unit of distance for the given pair. The result depends on the difference in positions. Shifting both positions together rotates the arrows but preserves their dot product.

An original example uses q=k=(1,0)q=k=(1,0) and an illustrative θ=30∘\theta=30^\circ. At distance 1, the dot product is about 0.866; at 6, it is −1; at 12, it is 1 again. This is one pair and a deliberately simple frequency, not a model's full recipe. The original paper uses θi=10000−2i/d\theta_i=10000^{-2i/d} radians for pair index i=0,…,d/2−1i=0,\ldots,d/2-1.

The displayed result is a component of Scaled Dot-Product Attention, before scaling and Softmax. It is not a probability. In this variant, we rotate Q and K; V does not require this rotation. RoPE does not replace Causal Masking.

A single pair's return to the same orientation also shows why one must not assume every score decreases monotonically with distance. The relative-position property alone does not demonstrate that the entire model works correctly on arbitrarily long text.

I use AI-generated content as part of my daily learning process.