Skip to content
Road to Intelligence

Concept · Chapter 11: Inside Modern LLMs

Rotary Position Embedding (RoPE)

Should knowUnderstand12 minDifficulty

RoPE encodes a token's position by rotating each pair of coordinates in its query and key by an angle proportional to the position, so the dot product between a query and a key depends only on how far apart the two tokens are.

The problem

Adding a position vector to each token's embedding mixes 'where' into 'what', and attention scores don't directly see the distance between two tokens.

The solution

Leave the embeddings alone and rotate the queries and keys instead: position m turns each 2-D pair of coordinates by m × θᵢ, with a different frequency θᵢ for each pair. Rotating both sides makes the score depend on the offset.

The consequence

RoPE became the standard in open LLMs. Its frequencies also became the main lever for longer contexts: interpolate positions, or raise the base so angles turn more slowly.

You should understand first

  1. Vectors
  2. Dot Product
  3. Embeddings
  4. Attention
  5. Probability and Distributions
  6. Softmax
  7. Self-Attention
  8. Positional Encoding
  9. Rotary Position Embedding (RoPE)

Position as an angle

Chapter 7's positional encoding added a position vector to each token. RoPE does something else. Su and colleagues encode absolute position with a rotation matrix applied in self-attention, which makes the attention score depend explicitly on the relative position of the two tokens Established.

Split a query's coordinates into pairs, and treat each pair as a point in a plane. At position mm, rotate the pair by the angle mθm\theta. Do the same to the key at position nn. Rotations preserve lengths, and the dot product of two rotated vectors depends only on the angle between them:

⟨R(mθ) q,  R(nθ) k⟩=⟨q,  R((n−m)θ) k⟩.\langle R(m\theta)\,q,\; R(n\theta)\,k \rangle = \langle q,\; R\big((n-m)\theta\big)\,k \rangle .

A tiny example

Take q=k=(1,0)q = k = (1, 0) and θ=30°\theta = 30°.

  • Query at position 2, key at position 5: rotated by 60° and 150°. They are 90° apart: score cos⁡90°=0\cos 90° = 0.
  • Query at 10, key at 13: rotated by 300° and 390° (= 30°). Again 90° apart: score 0.
  • Query at 4, key at 5: 30° apart: score cos⁡30°≈0.87\cos 30° \approx 0.87.

Same distance, same score, wherever the pair sits in the sequence.

Many clocks

A real head has many pairs, each turning at its own speed. RoFormer sets the frequencies to θi=10000−2(i−1)/d\theta_i = 10000^{-2(i-1)/d} Established, the same geometric spread as the original sine waves: fast pairs distinguish neighbours, slow pairs distinguish distant tokens. Llama 3 uses RoPE with a base of 500,000 Established; a larger base makes the slow pairs turn more slowly, which helps when positions run into the tens of thousands.

Going past the trained length

A model trained on 4,096 positions has never seen the angles of position 20,000. Chen and colleagues found that extrapolating RoPE beyond the trained length can produce catastrophically high attention scores; instead they linearly scale down the position indices to fit the original window, and extended LLaMA models to 32,768 tokens with fine-tuning of up to 1,000 steps Established. Llama 3 supports contexts of up to 128K tokens Established.

What to remember

  • Rotate q and k by angles proportional to their positions; leave values alone.
  • Rotating both by m and n: the score depends on n − m only.
  • Each pair of dimensions turns at its own frequency θᵢ = base^(−2(i−1)/d), base 10,000 in RoFormer.
  • Llama 3 uses a base of 500,000 and contexts up to 128K tokens.
  • Position interpolation: squeeze long positions into the trained range, then fine-tune briefly.

Key papers

Important

RoFormer: Enhanced Transformer with Rotary Position Embedding

Jianlin Su, Yu Lu et al. · 2021

Rotary position embedding (RoPE) is how most open LLMs encode word order: rotate queries and keys by position-dependent angles so attention scores depend on relative distance.

How to read it: Section 3.2.1 has the 2-D case that makes the idea clear; the general form is the same rotation applied to every pair of dimensions.

~35 min readarXiv:2104.09864✓ verified 2026-10-05

Watch