Skip to content
Road to Intelligence

Concept · Chapter 14: Reasoning Models

Group Relative Policy Optimization (GRPO)

FrontierKnow well11 minDifficulty

GRPO is a PPO-style policy-optimization method that judges each sampled response against the other responses to the same prompt, replacing PPO's learned value model with a group average.

The problem

PPO estimates how much better than expected each response was using a value model roughly as large as the policy, which doubles memory and adds a second network to train.

The solution

Sample a group of responses per prompt, score them, and use each response's reward relative to the group's mean (scaled by the group's spread) as its advantage.

The consequence

RL on long reasoning traces became cheaper to run, and GRPO became a standard choice for RL with verifiable rewards, with the side effect that groups where every answer scores the same teach nothing.

Better than what?

A reward of 1 says a response passed. To learn, the policy needs to know whether it did better than expected: that is the advantage. Policy-gradient methods push up responses with positive advantage and push down those with negative advantage.

PPO estimates “expected” with a learned value model (a critic) that predicts the reward from the prompt and partial response. For a large language model the critic is usually another network of comparable size, trained alongside the policy.

DeepSeekMath introduced Group Relative Policy Optimization, a PPO variant that drops the critic and instead estimates the baseline from the scores of a group of responses sampled for the same prompt, reducing memory use. Established

The group baseline

For one prompt, sample GG responses and score them r1,…,rGr_1, \dots, r_G. Each response's advantage is

Ai=ri−mean⁡(r)std⁡(r)+ϵ.A_i = \frac{r_i - \operatorname{mean}(r)}{\operatorname{std}(r) + \epsilon}.

Tiny example. Four responses with rewards [1,0,0,1][1, 0, 0, 1]. The mean is 0.5 and the population standard deviation is 0.5, so the advantages are [1,−1,−1,1][1, -1, -1, 1] (ignoring the tiny ϵ\epsilon). The two passing responses are pushed up equally and the two failing ones down.

What it does. The group plays the role of the critic: “expected” means “what this model typically scores on this prompt right now”. It is the same idea as comparing a pipeline run against its own recent runs instead of against a separately trained forecast.

Groups with nothing to say

If all four responses pass, r=[1,1,1,1]r = [1, 1, 1, 1]: every ri−mean⁡(r)r_i - \operatorname{mean}(r) is zero and the standard deviation is zero too, so implementations must guard the division and the prompt contributes no gradient. The same happens if all four fail. Problems the model always solves, or never solves, waste compute.

So the curriculum matters: useful training problems sit where the model sometimes succeeds. This is the RL version of the observation in the consensus lab that more samples help only when the generator has a chance.

The rest of the objective

The advantage is one ingredient. GRPO, like PPO, also uses:

  • a probability ratio between the current policy and the one that generated the samples, so several update steps can reuse one batch;
  • clipping of that ratio, so one batch cannot move the policy too far;
  • a KL penalty toward a reference model, discouraging drift into degenerate text;
  • choices about averaging over tokens, which decide how much each token of a long or short response counts.

Reproducing a GRPO result requires all of these details, and published implementations differ.

The DeepSeek-R1 report used GRPO for its large-scale RL stages. Established That connection, an open model trained with a documented, critic-free method, is much of why GRPO spread quickly.

Mini experiment

Compute the advantages by hand for rewards [1,1,1,0][1, 1, 1, 0] and [0,0,0,1][0, 0, 0, 1]. Which single response gets the larger push in each case, and why does a rare success get a bigger advantage than a common one? Then explain what happens to a prompt where all eight samples score 1.

What to remember

  • Advantage of response i ≈ (rᵢ − mean of the group) ÷ (std of the group + ε).
  • Rewards [1, 0, 0, 1] give advantages close to [1, −1, −1, 1].
  • No separate value (critic) model: the group is the baseline.
  • All-correct or all-wrong groups give zero advantage; pick problems at the edge of the model's ability.
  • The advantage formula is not the whole objective: ratios, clipping and a KL term still matter.

Key papers

Important

DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

DeepSeek-AI et al. · 2025

An openly released reasoning model, with a detailed account of training long chains of reasoning mainly through reinforcement learning on verifiable problems.

How to read it: Read the R1-Zero and R1 sections separately: the first is RL straight from a base model, the second a multi-stage recipe with cold-start data, and the distillation results are a third story.

~1 h readarXiv:2501.12948✓ verified 2026-09-26
Essential

Proximal Policy Optimization Algorithms

John Schulman, Filip Wolski et al. · 2017

PPO: a simple, robust actor–critic policy-gradient method. It became the default RL algorithm in many labs and was the optimiser in InstructGPT-style RLHF.

~30 min readarXiv:1707.06347✓ verified 2026-09-26
Important

DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Zhihong Shao, Peiyi Wang et al. · 2024

Introduced GRPO, the group-relative RL method later used to train DeepSeek-R1 and widely adopted for RL with verifiable rewards.

How to read it: For GRPO, go to the RL section and compare its objective with PPO's term by term; the data pipeline sections are a separate, also useful, story.

~1 h readarXiv:2402.03300✓ verified 2026-10-06