Skip to content
Road to Intelligence

Concept · Chapter 14: Reasoning Models

Verifiers and Best-of-N

Must knowKnow well12 minDifficulty

A verifier scores or checks candidate solutions so a system can keep the best one, and the quality of that check, not the number of candidates, often decides the final accuracy.

The problem

Sampling many solutions makes a correct one likely to exist, but someone still has to recognise it among the wrong ones.

The solution

Score each candidate with a check: an exact answer comparison, tests, a proof checker, or a learned model trained to judge correctness, and keep the highest-scoring candidate (best-of-N).

The consequence

Generation and selection become separate problems with separate fixes, and every selector has blind spots that more sampling can exploit.

Two failures with different fixes

Generate a batch of candidate solutions. Exactly one of two things has gone wrong if the system's final answer is wrong:

  • Generation failure: no candidate was correct. More samples help only if the generator sometimes succeeds; otherwise you need a better generator.
  • Selection failure: a correct candidate existed and was discarded. You need a better selector.

Measuring both requires a known correct answer, which is why benchmarks report an “oracle” or pass@N number (some candidate was right) alongside the selected accuracy. The gap between them is the selector's cost.

Kinds of check

CheckExampleWhat it misses
Exact answer matchFinal number equals 21Right answer, wrong reasoning
Unit testsProgram passes 12 testsInputs the tests never try
Formal checkerProof assistant accepts a proofWhether the theorem stated is the one you meant
Learned outcome scorer (ORM)Model trained to predict correct final answersAnything outside its training distribution
Learned process scorer (PRM)Model trained to judge each stepAnnotation errors; ambiguous steps

Only the first three are mechanical. A learned scorer is a model with its own errors, and its score is not a calibrated probability unless you have checked it.

Best-of-N in numbers

Cobbe and colleagues introduced GSM8K, 8.5K grade-school math word problems, and trained verifiers to judge model solutions; generating many candidates and keeping the one ranked highest by the verifier significantly improved performance, and scaled better with more data than fine-tuning alone. Established

In OpenAI's o1 announcement, AIME 2024 accuracy was reported as 74% with one sample, 83% with a 64-sample vote, and 93% when re-ranking 1,000 samples with a learned scoring function. Established A strong scorer extracted more from many samples than a vote did.

Tiny example (lab model). The consensus lab gives correct answers a base score of 0.75 and wrong answers 0.25, then adds independent noise in log-odds. With no noise the verifier finds a correct candidate whenever one exists. At the default settings (45% fresh accuracy, 16 candidates) the vote returns 21 only about half the time, because the commonest mistake competes with it, while the verifier's pick is correct in nearly every batch.

When more candidates hurt

The lab's noise is independent for every candidate, so more candidates mostly help. Real scorers are wrong in systematic ways: a learned judge may prefer confident tone or a particular wrong method. Then sampling more gives the scorer more chances to find a wrong answer it over-rates.

Gao and colleagues trained a proxy reward model against a fixed “gold” reward model and optimized against the proxy with best-of-N sampling and with RL; in both cases the gold score eventually fell as optimization against the proxy increased. Established This is Goodhart's law measured: once the score becomes the target, it stops tracking quality. Leave headroom, and evaluate the selected outputs with a check the selector never saw.

Where it shows up

Coding agents run tests before claiming success; that is a verifier gating output. RL for reasoning uses the same checks as rewards, see verifiable rewards. And process supervision moves the check from the final answer to each step.

Mini experiment

In the lab, set fresh accuracy to 20%, noise to 100% and candidates to 4. Play ten seeds, tallying how often a correct answer exists and how often the verifier picks it. Repeat with 32 candidates. Then explain, in two sentences, why independent noise behaves differently from a scorer that systematically likes the answer 42.

Why should I care?

As a researcher

Verifiers turn sampling into accuracy, supply the rewards for RL on reasoning, and their failure modes (overoptimization, blind spots) set the limits of both.

As an engineer

Whenever you can write a check (tests, a schema, a calculator), you can generate several candidates and keep one that passes, which is often cheaper than a bigger model.

Modern systems that depend on it

  • RL with verifiable rewards
  • process reward models
  • test-time compute allocation
  • coding agents that run tests

Historical context

Before

A system returned whatever the model generated first, so its accuracy was the model's single-sample accuracy.

After

Systems generate several candidates and keep the one a verifier prefers, and they report a generation failure separately from a selection failure.

Used today

Coding tools run tests on generated patches, math systems check answers, and learned reward models rerank samples and supply training rewards.

What to remember

  • Generation failure: no candidate is correct. Selection failure: one is, but the selector picked another.
  • Best-of-N keeps the highest-scored candidate; self-consistency keeps the most common answer.
  • Cobbe et al. (2021) introduced GSM8K and showed trained verifiers plus many samples improved accuracy.
  • Checks differ in kind: exact match, tests, proof checkers and learned scorers each have different blind spots.
  • Pushing best-of-N hard against an imperfect scorer can lower true quality (Gao et al., 2022).
  • Oracle selection is an upper bound for analysis, not a deployable strategy.

Key papers

Important

Scaling Laws for Reward Model Overoptimization

Leo Gao, John Schulman, Jacob Hilton · 2022

Explains why maximizing a learned judge can diverge from the intended outcome.

How to read it: Read the synthetic-gold-reward setup before interpreting the scaling curves.

~30 min readarXiv:2210.10760✓ verified 2026-10-04
Important

Training Verifiers to Solve Math Word Problems

Karl Cobbe, Vineet Kosaraju et al. · 2021

Introduced GSM8K, the grade-school math benchmark used for years afterwards, and the generate-many-then-verify recipe that best-of-N selection still follows.

How to read it: Read the sections on the verifier and on how test performance changes with the number of sampled completions; that trade-off is Chapter 14's selection problem.

~35 min readarXiv:2110.14168✓ verified 2026-10-06
Important

Let's Verify Step by Step

Hunter Lightman, Vineet Kosaraju et al. · 2023

The clearest comparison of rewarding steps versus rewarding final answers, and the source of PRM800K, a widely used set of step-level labels.

How to read it: Both reward models are used to pick the best of many samples (best-of-N), not to train the generator with RL; keep that setup in mind when generalising.

~45 min readarXiv:2305.20050✓ verified 2026-10-06