Concept · Chapter 11: Inside Modern LLMs
Rotary Position Embedding (RoPE)
RoPE encodes a token's position by rotating each pair of coordinates in its query and key by an angle proportional to the position, so the dot product between a query and a key depends only on how far apart the two tokens are.
The problem
Adding a position vector to each token's embedding mixes 'where' into 'what', and attention scores don't directly see the distance between two tokens.
The solution
Leave the embeddings alone and rotate the queries and keys instead: position m turns each 2-D pair of coordinates by m × θᵢ, with a different frequency θᵢ for each pair. Rotating both sides makes the score depend on the offset.
The consequence
RoPE became the standard in open LLMs. Its frequencies also became the main lever for longer contexts: interpolate positions, or raise the base so angles turn more slowly.
You should understand first
- Vectors
- Dot Product
- Embeddings
- Attention
- Probability and Distributions
- Softmax
- Self-Attention
- Positional Encoding
- Rotary Position Embedding (RoPE)
Position as an angle
Chapter 7's positional encoding added a position vector to each token. RoPE does something else. Su and colleagues encode absolute position with a rotation matrix applied in self-attention, which makes the attention score depend explicitly on the relative position of the two tokens Established.
Split a query's coordinates into pairs, and treat each pair as a point in a plane. At position , rotate the pair by the angle . Do the same to the key at position . Rotations preserve lengths, and the dot product of two rotated vectors depends only on the angle between them:
A tiny example
Take and .
- Query at position 2, key at position 5: rotated by 60° and 150°. They are 90° apart: score .
- Query at 10, key at 13: rotated by 300° and 390° (= 30°). Again 90° apart: score 0.
- Query at 4, key at 5: 30° apart: score .
Same distance, same score, wherever the pair sits in the sequence.
Many clocks
A real head has many pairs, each turning at its own speed. RoFormer sets the frequencies to Established, the same geometric spread as the original sine waves: fast pairs distinguish neighbours, slow pairs distinguish distant tokens. Llama 3 uses RoPE with a base of 500,000 Established; a larger base makes the slow pairs turn more slowly, which helps when positions run into the tens of thousands.
Going past the trained length
A model trained on 4,096 positions has never seen the angles of position 20,000. Chen and colleagues found that extrapolating RoPE beyond the trained length can produce catastrophically high attention scores; instead they linearly scale down the position indices to fit the original window, and extended LLaMA models to 32,768 tokens with fine-tuning of up to 1,000 steps Established. Llama 3 supports contexts of up to 128K tokens Established.
What to remember
- Rotate q and k by angles proportional to their positions; leave values alone.
- Rotating both by m and n: the score depends on n − m only.
- Each pair of dimensions turns at its own frequency θᵢ = base^(−2(i−1)/d), base 10,000 in RoFormer.
- Llama 3 uses a base of 500,000 and contexts up to 128K tokens.
- Position interpolation: squeeze long positions into the trained range, then fine-tune briefly.
Key papers
RoFormer: Enhanced Transformer with Rotary Position Embedding
Jianlin Su, Yu Lu et al. · 2021
Rotary position embedding (RoPE) is how most open LLMs encode word order: rotate queries and keys by position-dependent angles so attention scores depend on relative distance.
How to read it: Section 3.2.1 has the 2-D case that makes the idea clear; the general form is the same rotation applied to every pair of dimensions.
Extending Context Window of Large Language Models via Positional Interpolation
Shouyuan Chen, Sherman Wong et al. · 2023
A simple trick for using RoPE models beyond their trained length: squeeze new positions into the old range instead of extrapolating.
Watch
Stanford Online
Stanford CS336 Lang. Modeling from Scratch | Spring 2025 | Lec. 3: Architectures, Hyperparameters
A survey of what changed inside the Transformer block between 2017 and today's open models, and which choices nearly everyone now agrees on.