Skip to content
Road to Intelligence

Concept · Chapter 14: Reasoning Models

RL with Verifiable Rewards

Must knowKnow well12 minDifficulty

Reinforcement learning with verifiable rewards trains a model on problems whose results a program can check, such as math answers or code tests, and raises the probability of the responses that pass.

The problem

Preference-based RL needs human judgments or a learned reward model, both expensive and exploitable, and supervised traces only teach a model to imitate solutions someone already wrote.

The solution

Use tasks with mechanical checks, sample responses from the current model, score each with the check, and update the policy toward responses that scored above average.

The consequence

Models learned to produce long, self-checking chains of thought that improved math and coding results, but only where checks exist, and every gap in a check is a behaviour the model can learn.

The checker supplies the signal

An arithmetic answer can be compared with a reference. A program can be run against tests. Each check turns a task into an environment that hands out rewards with no human in the loop. Reinforcement learning then does what it did in Chapter 5: try, score, and make high-scoring behaviour more likely.

Reinforcement learning with verifiable rewards uses an executable or rule-based check, such as agreement with a known answer or passing specified tests, to produce training feedback. Established

Compare the alternatives from Chapter 10. RLHF learns a reward model from human preferences, which can judge open-ended answers but is itself a model with exploitable errors. Supervised fine-tuning imitates solutions someone wrote, so it cannot exceed them. A verifiable reward is narrower than either, but cheap, consistent and not a neural network that can be flattered.

The loop

  1. Sample: for each training problem, generate one or more complete responses from the current model.
  2. Check: compute each response's reward (1 if the answer matches, perhaps a small reward for the right format).
  3. Compare: estimate how much better or worse each response is than expected, for example against the group average (GRPO).
  4. Update: change the weights to make better-than-expected responses more likely, with limits on how far one step can move the policy.

Inference-time selection (best-of-N) uses the same check but stops after step 2. RL is what makes the next attempt better.

The exact gradient, in a tiny case

The reward lab has four fixed responses with logits θi\theta_i, probabilities pi=softmax⁡(θ)ip_i = \operatorname{softmax}(\theta)_i and rewards rir_i. Expected reward is J=∑ipiriJ = \sum_i p_i r_i, and

∂J∂θi=pi (ri−J).\frac{\partial J}{\partial \theta_i} = p_i\,(r_i - J).

What it does. A response gains logit in proportion to how likely it already is and how much better than average it scored. Numbers. Start uniform, pi=0.25p_i = 0.25. With outcome reward r=[1,1,0,0]r = [1, 1, 0, 0], J=0.5J = 0.5, so each correct response's logit rises by 0.25×0.5=0.1250.25 \times 0.5 = 0.125 and each wrong one's falls by 0.1250.125. After one update the correct responses each have probability 0.281.

Real training cannot compute JJ exactly: it estimates the gradient from sampled token sequences, which is noisy, and handles thousands of tokens per response. The lab removes that noise on purpose to isolate one question: what does this reward make more likely?

What R1 reported

The January 2025 DeepSeek-R1 report described R1-Zero, trained with large-scale RL using rule-based rewards directly from a base model without supervised fine-tuning first, and R1, which added cold-start data and further stages; it also described distilling R1's reasoning into smaller models. Established The report attributed behaviours such as self-verification and longer reasoning to the RL stage. “RL alone” still begins from a large pretrained model, which supplies the language and much of the mathematics.

How much RL with verifiable rewards creates new problem-solving ability, versus making the base model's existing good samples more likely, is debated and depends on model, task and evaluation budget. Active research One useful comparison: the base model's accuracy when allowed many samples versus the trained model's accuracy with one.

What a check cannot see

  • Incomplete tests reward code that passes them without meeting the specification.
  • Answer-only checks reward lucky or memorised answers; see process supervision.
  • Exploitable checkers: if the model can affect the checker (edit the test file, print the expected string), it may learn to.
  • Coverage: tasks without checks, such as writing, research design or judgement calls, get no reward at all.

Keep the training checks separate from the evaluation set, and inspect samples, not only the reward curve.

Mini experiment

In the lab, train with Correct final answer for 60 updates and note the valid-derivation probability. Reset, train with Valid steps, and note it again. Then write down a reward for a coding task that would let a model score well without writing a correct program, and one extra check that would close the gap.

Why should I care?

As a researcher

Checkable rewards made RL on long reasoning traces practical at scale, and the open question of what RL adds beyond the base model's own samples is central to current work.

As an engineer

If your domain has a reliable automatic check, it can serve as a training reward as well as an evaluation, and its loopholes will be found.

Modern systems that depend on it

  • GRPO
  • reasoning models such as DeepSeek-R1
  • distilled reasoning models
  • coding models trained against tests

Historical context

Before

Reasoning behaviour came from prompting or from imitating human-written solutions, and RL for language models relied on learned preference rewards.

After

Models are trained by trial and error against programmatic checks, and longer, self-correcting reasoning emerges where it raises the reward.

Used today

Recent reasoning models are reported to rely on RL against math answers, code tests and other rule-based checks.

What to remember

  • Verifiable means a program computes the reward, not that the reward is complete.
  • Sample → check → compare with the average → update weights; inference-time selection skips the update.
  • Exact gradient of expected reward for a softmax policy: ∂J/∂θᵢ = pᵢ(rᵢ − J).
  • Outcome-only rewards cannot separate valid from lucky reasoning.
  • DeepSeek-R1-Zero applied RL to a base model with no supervised fine-tuning first; R1 used a multi-stage recipe.
  • Problems where every sample passes, or every sample fails, give no learning signal.

Key papers

Important

DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

DeepSeek-AI et al. · 2025

An openly released reasoning model, with a detailed account of training long chains of reasoning mainly through reinforcement learning on verifiable problems.

How to read it: Read the R1-Zero and R1 sections separately: the first is RL straight from a base model, the second a multi-stage recipe with cold-start data, and the distillation results are a third story.

~1 h readarXiv:2501.12948✓ verified 2026-09-26
Important

STaR: Bootstrapping Reasoning With Reasoning

Eric Zelikman, Yuhuai Wu et al. · 2022 · NeurIPS 2022

An early, clear version of the loop behind synthetic reasoning data: a model generates rationales, keeps the ones that reach correct answers, and trains on them.

How to read it: Notice the filter is the final answer only, so a rationale that reaches the right answer by a flawed route is kept; compare Chapter 14's lucky answer.

~35 min readarXiv:2203.14465✓ verified 2026-10-06
Important

DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Zhihong Shao, Peiyi Wang et al. · 2024

Introduced GRPO, the group-relative RL method later used to train DeepSeek-R1 and widely adopted for RL with verifiable rewards.

How to read it: For GRPO, go to the RL section and compare its objective with PPO's term by term; the data pipeline sections are a separate, also useful, story.

~1 h readarXiv:2402.03300✓ verified 2026-10-06

Watch

3 h 31 min

Andrej Karpathy

Deep Dive into LLMs like ChatGPT

A long, general-audience walk through the whole pipeline behind a chat model, from internet text to tokens to pretraining to post-training.

Should know