Skip to content
Road to Intelligence

Concept · Chapter 14: Reasoning Models

Process vs Outcome Supervision

Should knowKnow well10 minDifficulty

Outcome supervision gives feedback on a solution's final result; process supervision gives feedback on each intermediate step, so a lucky answer reached by invalid steps is not rewarded.

The problem

A final-answer label cannot say where a derivation went wrong, and it treats “6 × 7 = 40, 40 ÷ 2 = 21” exactly like the valid derivation because both end in 21.

The solution

Label or score every step, and train a process reward model that predicts whether each step is correct; use it to select solutions or as a training signal.

The consequence

Feedback becomes denser and more precise about errors, at the price of expensive, sometimes ambiguous step labels and a learned checker that can itself be wrong.

Two places for a label

Compare two solutions to 6×7÷26 \times 7 \div 2:

  • Valid: 6×7=426 \times 7 = 42, then 42÷2=2142 \div 2 = 21.
  • Lucky: 6×7=406 \times 7 = 40, then 40÷2=2140 \div 2 = 21. Two wrong steps that happen to cancel.

An outcome label looks at the last line and marks both correct. A process label looks at each step and rejects the lucky solution at its first line. If you train on outcome labels, the lucky solution is as good as the valid one; the model has no reason to prefer valid steps when they lead to the same number.

In data-engineering terms: outcome supervision checks that the final report's total matches; process supervision checks every transformation in the pipeline. A matching total can hide two bugs that cancel.

The evidence

Lightman and colleagues compared outcome and process supervision for training reward models on the MATH dataset and found process supervision significantly better; their process-supervised model solved 78% of a representative subset of MATH test problems, and they released PRM800K, 800,000 step-level human labels. Established

Read the setup before generalising. The reward models were judged by how well they picked the best of many sampled solutions; the paper compares selectors, not RL training recipes. Active learning (choosing which solutions to label) also made the labelling budget go further.

The price of steps

  • Labels are expensive. Someone, or something, must judge every step, and long solutions have many.
  • Steps are ambiguous. Is “Let x be the number of trays” correct? Is an unnecessary but true step good? Labellers disagree.
  • A learned PRM is a model. It generalises imperfectly and can be fooled, just like an outcome verifier.
  • Valid is not the same as useful. A long chain of true but irrelevant steps passes a step checker.

Where steps are mechanically checkable (arithmetic, formal proofs, code that runs), process checks can be programs instead of models, which removes the labelling cost and most of the ambiguity.

Surface features are not process

“More detailed” sounds like “better process”, but detail can be rewarded by length alone. In the reward lab, choose More words: the verbose mistake, which repeats a wrong step with confident filler, climbs to about 97% probability within 60 updates, and the probability of a correct answer falls to about 2%. Choose Valid steps and the valid derivation reaches about 99%. Choose Correct final answer and the valid and lucky derivations rise together and stay exactly tied, each near 49%.

Mini experiment

Write three solutions to a two-step problem of your own: one valid, one with compensating errors, one verbose and wrong. Score each with three rules: final answer, every step valid, and number of lines. Which rule would you trust as a training reward, and which could a model satisfy without solving anything?

What to remember

  • Outcome supervision: one label for the result. Process supervision: a label for every step.
  • Final-answer rewards cannot tell a valid derivation from a lucky one.
  • Lightman et al. found process supervision outperformed outcome supervision on MATH and released 800,000 step labels (PRM800K).
  • Their comparison selected among sampled solutions (best-of-N); it was not an RL training run.
  • Rewarding surface features such as length teaches the surface feature, as the reward lab shows.

Key papers

Important

Training Verifiers to Solve Math Word Problems

Karl Cobbe, Vineet Kosaraju et al. · 2021

Introduced GSM8K, the grade-school math benchmark used for years afterwards, and the generate-many-then-verify recipe that best-of-N selection still follows.

How to read it: Read the sections on the verifier and on how test performance changes with the number of sampled completions; that trade-off is Chapter 14's selection problem.

~35 min readarXiv:2110.14168✓ verified 2026-10-06
Important

Let's Verify Step by Step

Hunter Lightman, Vineet Kosaraju et al. · 2023

The clearest comparison of rewarding steps versus rewarding final answers, and the source of PRM800K, a widely used set of step-level labels.

How to read it: Both reward models are used to pick the best of many samples (best-of-N), not to train the generator with RL; keep that setup in mind when generalising.

~45 min readarXiv:2305.20050✓ verified 2026-10-06