Concept · Chapter 14: Reasoning Models
Group Relative Policy Optimization (GRPO)
GRPO is a PPO-style policy-optimization method that judges each sampled response against the other responses to the same prompt, replacing PPO's learned value model with a group average.
The problem
PPO estimates how much better than expected each response was using a value model roughly as large as the policy, which doubles memory and adds a second network to train.
The solution
Sample a group of responses per prompt, score them, and use each response's reward relative to the group's mean (scaled by the group's spread) as its advantage.
The consequence
RL on long reasoning traces became cheaper to run, and GRPO became a standard choice for RL with verifiable rewards, with the side effect that groups where every answer scores the same teach nothing.
You should understand first
- Text as Data
- Probability and Distributions
- Conditional Probability and Bayes' Theorem
- Probability of Sequences
- Language Modeling
- Vectors
- Dot Product
- Embeddings
- Attention
- Softmax
- Self-Attention
- Multi-Head Attention
- Causal Masking
- Positional Encoding
- Residual Connections
- Layer Normalization
- Feed-Forward Sublayer (MLP)
- The Transformer Block
- Encoder, Decoder & Encoder–Decoder
- Pretrain, Then Fine-Tune
- Entropy
- Loss Functions
- Cross-Entropy Loss
- Supervised Fine-Tuning
- The Turing Test
- Symbolic AI
- Logic and Rules
- Expert Systems
- Knowledge Representation
- The Knowledge-Acquisition Bottleneck
- From Rules to Learning
- Supervised, Unsupervised and Self-Supervised Learning
- Autoregressive Next-Token Prediction
- Decoding: Greedy, Temperature, Top-k, Top-p
- Learning from Comparisons
- Features, Labels and Tasks
- Derivatives and Gradients
- Gradient Descent
- Linear Regression
- Logistic Regression
- Reward Models
- Expected Value and Variance
- Reinforcement Learning
- MDPs, Policies and Value
- Policy Gradients
- KL Divergence
- Reinforcement Learning from Human Feedback
- Q-Learning
- Actor–Critic Methods
- PPO for Language Models
- Pretraining at Scale
- GPT-1 → GPT-2 → GPT-3
- In-Context Learning
- Chain of Thought
- Sampling and Uncertainty
- Self-Consistency
- Verifiers and Best-of-N
- RL with Verifiable Rewards
- Group Relative Policy Optimization (GRPO)
Better than what?
A reward of 1 says a response passed. To learn, the policy needs to know whether it did better than expected: that is the advantage. Policy-gradient methods push up responses with positive advantage and push down those with negative advantage.
PPO estimates “expected” with a learned value model (a critic) that predicts the reward from the prompt and partial response. For a large language model the critic is usually another network of comparable size, trained alongside the policy.
DeepSeekMath introduced Group Relative Policy Optimization, a PPO variant that drops the critic and instead estimates the baseline from the scores of a group of responses sampled for the same prompt, reducing memory use. EstablishedThe group baseline
For one prompt, sample responses and score them . Each response's advantage is
Tiny example. Four responses with rewards . The mean is 0.5 and the population standard deviation is 0.5, so the advantages are (ignoring the tiny ). The two passing responses are pushed up equally and the two failing ones down.
What it does. The group plays the role of the critic: “expected” means “what this model typically scores on this prompt right now”. It is the same idea as comparing a pipeline run against its own recent runs instead of against a separately trained forecast.
Groups with nothing to say
If all four responses pass, : every is zero and the standard deviation is zero too, so implementations must guard the division and the prompt contributes no gradient. The same happens if all four fail. Problems the model always solves, or never solves, waste compute.
So the curriculum matters: useful training problems sit where the model sometimes succeeds. This is the RL version of the observation in the consensus lab that more samples help only when the generator has a chance.
The rest of the objective
The advantage is one ingredient. GRPO, like PPO, also uses:
- a probability ratio between the current policy and the one that generated the samples, so several update steps can reuse one batch;
- clipping of that ratio, so one batch cannot move the policy too far;
- a KL penalty toward a reference model, discouraging drift into degenerate text;
- choices about averaging over tokens, which decide how much each token of a long or short response counts.
Reproducing a GRPO result requires all of these details, and published implementations differ.
The DeepSeek-R1 report used GRPO for its large-scale RL stages. Established That connection, an open model trained with a documented, critic-free method, is much of why GRPO spread quickly.
Mini experiment
Compute the advantages by hand for rewards and . Which single response gets the larger push in each case, and why does a rare success get a bigger advantage than a common one? Then explain what happens to a prompt where all eight samples score 1.
What to remember
- Advantage of response i ≈ (rᵢ − mean of the group) ÷ (std of the group + ε).
- Rewards [1, 0, 0, 1] give advantages close to [1, −1, −1, 1].
- No separate value (critic) model: the group is the baseline.
- All-correct or all-wrong groups give zero advantage; pick problems at the edge of the model's ability.
- The advantage formula is not the whole objective: ratios, clipping and a KL term still matter.
Key papers
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
DeepSeek-AI et al. · 2025
An openly released reasoning model, with a detailed account of training long chains of reasoning mainly through reinforcement learning on verifiable problems.
How to read it: Read the R1-Zero and R1 sections separately: the first is RL straight from a base model, the second a multi-stage recipe with cold-start data, and the distillation results are a third story.
Proximal Policy Optimization Algorithms
John Schulman, Filip Wolski et al. · 2017
PPO: a simple, robust actor–critic policy-gradient method. It became the default RL algorithm in many labs and was the optimiser in InstructGPT-style RLHF.
DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Zhihong Shao, Peiyi Wang et al. · 2024
Introduced GRPO, the group-relative RL method later used to train DeepSeek-R1 and widely adopted for RL with verifiable rewards.
How to read it: For GRPO, go to the RL section and compare its objective with PPO's term by term; the data pipeline sections are a separate, also useful, story.