Concept · Chapter 14: Reasoning Models
Verifiers and Best-of-N
A verifier scores or checks candidate solutions so a system can keep the best one, and the quality of that check, not the number of candidates, often decides the final accuracy.
The problem
Sampling many solutions makes a correct one likely to exist, but someone still has to recognise it among the wrong ones.
The solution
Score each candidate with a check: an exact answer comparison, tests, a proof checker, or a learned model trained to judge correctness, and keep the highest-scoring candidate (best-of-N).
The consequence
Generation and selection become separate problems with separate fixes, and every selector has blind spots that more sampling can exploit.
You should understand first
- The Turing Test
- Symbolic AI
- Logic and Rules
- Expert Systems
- Knowledge Representation
- The Knowledge-Acquisition Bottleneck
- From Rules to Learning
- Supervised, Unsupervised and Self-Supervised Learning
- Text as Data
- Probability and Distributions
- Conditional Probability and Bayes' Theorem
- Probability of Sequences
- Language Modeling
- Vectors
- Dot Product
- Embeddings
- Attention
- Softmax
- Self-Attention
- Causal Masking
- Entropy
- Loss Functions
- Cross-Entropy Loss
- Autoregressive Next-Token Prediction
- Decoding: Greedy, Temperature, Top-k, Top-p
- Learning from Comparisons
- Features, Labels and Tasks
- Derivatives and Gradients
- Gradient Descent
- Linear Regression
- Logistic Regression
- Reward Models
- Pretraining at Scale
- GPT-1 → GPT-2 → GPT-3
- In-Context Learning
- Chain of Thought
- Expected Value and Variance
- Sampling and Uncertainty
- Self-Consistency
- Verifiers and Best-of-N
Two failures with different fixes
Generate a batch of candidate solutions. Exactly one of two things has gone wrong if the system's final answer is wrong:
- Generation failure: no candidate was correct. More samples help only if the generator sometimes succeeds; otherwise you need a better generator.
- Selection failure: a correct candidate existed and was discarded. You need a better selector.
Measuring both requires a known correct answer, which is why benchmarks report an “oracle” or pass@N number (some candidate was right) alongside the selected accuracy. The gap between them is the selector's cost.
Kinds of check
| Check | Example | What it misses |
|---|---|---|
| Exact answer match | Final number equals 21 | Right answer, wrong reasoning |
| Unit tests | Program passes 12 tests | Inputs the tests never try |
| Formal checker | Proof assistant accepts a proof | Whether the theorem stated is the one you meant |
| Learned outcome scorer (ORM) | Model trained to predict correct final answers | Anything outside its training distribution |
| Learned process scorer (PRM) | Model trained to judge each step | Annotation errors; ambiguous steps |
Only the first three are mechanical. A learned scorer is a model with its own errors, and its score is not a calibrated probability unless you have checked it.
Best-of-N in numbers
Cobbe and colleagues introduced GSM8K, 8.5K grade-school math word problems, and trained verifiers to judge model solutions; generating many candidates and keeping the one ranked highest by the verifier significantly improved performance, and scaled better with more data than fine-tuning alone. EstablishedIn OpenAI's o1 announcement, AIME 2024 accuracy was reported as 74% with one sample, 83% with a 64-sample vote, and 93% when re-ranking 1,000 samples with a learned scoring function. Established A strong scorer extracted more from many samples than a vote did.
Tiny example (lab model). The consensus lab gives correct answers a base score of 0.75 and wrong answers 0.25, then adds independent noise in log-odds. With no noise the verifier finds a correct candidate whenever one exists. At the default settings (45% fresh accuracy, 16 candidates) the vote returns 21 only about half the time, because the commonest mistake competes with it, while the verifier's pick is correct in nearly every batch.
When more candidates hurt
The lab's noise is independent for every candidate, so more candidates mostly help. Real scorers are wrong in systematic ways: a learned judge may prefer confident tone or a particular wrong method. Then sampling more gives the scorer more chances to find a wrong answer it over-rates.
Gao and colleagues trained a proxy reward model against a fixed “gold” reward model and optimized against the proxy with best-of-N sampling and with RL; in both cases the gold score eventually fell as optimization against the proxy increased. Established This is Goodhart's law measured: once the score becomes the target, it stops tracking quality. Leave headroom, and evaluate the selected outputs with a check the selector never saw.
Where it shows up
Coding agents run tests before claiming success; that is a verifier gating output. RL for reasoning uses the same checks as rewards, see verifiable rewards. And process supervision moves the check from the final answer to each step.
Mini experiment
In the lab, set fresh accuracy to 20%, noise to 100% and candidates to 4. Play ten seeds, tallying how often a correct answer exists and how often the verifier picks it. Repeat with 32 candidates. Then explain, in two sentences, why independent noise behaves differently from a scorer that systematically likes the answer 42.
Why should I care?
As a researcher
Verifiers turn sampling into accuracy, supply the rewards for RL on reasoning, and their failure modes (overoptimization, blind spots) set the limits of both.
As an engineer
Whenever you can write a check (tests, a schema, a calculator), you can generate several candidates and keep one that passes, which is often cheaper than a bigger model.
Modern systems that depend on it
- RL with verifiable rewards
- process reward models
- test-time compute allocation
- coding agents that run tests
Historical context
Before
A system returned whatever the model generated first, so its accuracy was the model's single-sample accuracy.
After
Systems generate several candidates and keep the one a verifier prefers, and they report a generation failure separately from a selection failure.
Used today
Coding tools run tests on generated patches, math systems check answers, and learned reward models rerank samples and supply training rewards.
What to remember
- Generation failure: no candidate is correct. Selection failure: one is, but the selector picked another.
- Best-of-N keeps the highest-scored candidate; self-consistency keeps the most common answer.
- Cobbe et al. (2021) introduced GSM8K and showed trained verifiers plus many samples improved accuracy.
- Checks differ in kind: exact match, tests, proof checkers and learned scorers each have different blind spots.
- Pushing best-of-N hard against an imperfect scorer can lower true quality (Gao et al., 2022).
- Oracle selection is an upper bound for analysis, not a deployable strategy.
Key papers
Scaling Laws for Reward Model Overoptimization
Leo Gao, John Schulman, Jacob Hilton · 2022
Explains why maximizing a learned judge can diverge from the intended outcome.
How to read it: Read the synthetic-gold-reward setup before interpreting the scaling curves.
Training Verifiers to Solve Math Word Problems
Karl Cobbe, Vineet Kosaraju et al. · 2021
Introduced GSM8K, the grade-school math benchmark used for years afterwards, and the generate-many-then-verify recipe that best-of-N selection still follows.
How to read it: Read the sections on the verifier and on how test performance changes with the number of sampled completions; that trade-off is Chapter 14's selection problem.
Let's Verify Step by Step
Hunter Lightman, Vineet Kosaraju et al. · 2023
The clearest comparison of rewarding steps versus rewarding final answers, and the source of PRM800K, a widely used set of step-level labels.
How to read it: Both reward models are used to pick the best of many samples (best-of-N), not to train the generator with RL; keep that setup in mind when generalising.