Concept · Chapter 14: Reasoning Models
Self-Consistency
Self-consistency samples several independent reasoning paths for the same question and returns the final answer they most often agree on.
The problem
One sampled or greedy chain of thought can go wrong at any step, and its single answer carries no sign of how reliable it is.
The solution
Sample many chains at a non-zero temperature, extract each final answer, and take a plurality vote; disagreement doubles as a signal of uncertainty.
The consequence
Accuracy rises with the number of samples on tasks with checkable final answers, at a proportional inference cost, but a vote amplifies whatever mistake the samples share.
You should understand first
- Text as Data
- Probability and Distributions
- Conditional Probability and Bayes' Theorem
- Probability of Sequences
- Language Modeling
- Vectors
- Dot Product
- Embeddings
- Attention
- Softmax
- Self-Attention
- Causal Masking
- Entropy
- Loss Functions
- Cross-Entropy Loss
- Autoregressive Next-Token Prediction
- Pretraining at Scale
- GPT-1 → GPT-2 → GPT-3
- In-Context Learning
- Chain of Thought
- Expected Value and Variance
- Sampling and Uncertainty
- Self-Consistency
Vote on answers, not on wording
Two chains can be worded completely differently and still end in 21. Self-consistency ignores the wording and counts final answers. If eight samples end in 21, 42, 21, 20, 21, 42, 21 and 21, the plurality answer is 21 with five votes.
Wang and colleagues proposed replacing greedy chain-of-thought decoding with sampling plus answer aggregation, reporting gains including +17.9% on GSM8K, +11.0% on SVAMP and +12.2% on AQuA. Established No training is involved; it is purely an inference-time procedure.
The intuition: a hard problem usually has several valid routes to one correct answer, while mistakes scatter across many different wrong answers. Correct paths agree with each other; wrong paths mostly disagree.
Availability is not selection
If each independent sample is correct with probability , the chance that at least one of is correct is
Tiny example. With and : . A correct answer is present 68% of the time. But the vote returns the most common answer, and with a single dominant wrong answer can easily be more common than the right one. Availability measures what a perfect selector could do; the vote is a specific, imperfect selector.
The same formula fails in a second way: it assumes every attempt is an independent coin flip with the same . Plugging in an average accuracy over a mix of easy and hard tasks overstates the gain, because on the hard tasks is near zero however many samples you draw.
When agreement misleads
In the consensus lab, the commonest mistake (forgetting to halve, giving 42) can outvote the correct 21 even when 21 appears several times. Push “Copy the first answer” to 100% and every candidate repeats the first one: thirty-two votes, one opinion.
Real models need no literal copying to fail together. Samples share a model, a prompt and a training history, so they can share a misreading of the question. A unanimous vote on a misread question is confident and wrong. The fix is diversity of evidence, not more samples from the same source: different prompts, tools, or an independent check.
Practical notes
- Cost. samples cost about times the generation tokens. Samples can run in parallel, so latency grows less than total compute does.
- Ties. Decide in advance. The lab abstains on a tie rather than picking arbitrarily.
- Answer extraction. Normalise answers before counting (“21”, “21.0” and “twenty-one” are one answer).
- Confidence. The vote share is a useful uncertainty signal: 15 of 16 agreeing is different from 6 of 16. It is not a calibrated probability.
- Free-form outputs such as essays or code have no exact answer to count; they need a scorer or tests, see verifiers.
OpenAI reported o1 at 74% on AIME 2024 with one sample per problem and 83% with a vote over 64 samples. Established Those are the company's own figures, and they illustrate the trade: about 64 times the samples for nine more points.
Mini experiment
In the lab, keep the default settings and press New seed ten times, playing each batch. Tally how often the vote returns 21, 42 or a tie, and how often the verifier's pick is 21. Then set the candidate budget to 32 and repeat. Which selector benefits more from more candidates at this accuracy?
What to remember
- Sample N chains, extract final answers, return the most common one.
- Wang et al. reported +17.9% on GSM8K over greedy chain-of-thought decoding.
- 1 − (1 − p)^N is the chance some sample is right, not the chance the vote is right.
- Shared mistakes win votes: correlated samples add cost without adding evidence.
- Voting needs answers that can be compared exactly; free-form text needs another selector.
Key papers
Self-Consistency Improves Chain of Thought Reasoning in Language Models
Xuezhi Wang, Jason Wei et al. · 2022 · ICLR 2023
Turned extra inference compute into accuracy without any training: sample many reasoning paths and keep the answer they most often reach.
How to read it: Look at how accuracy grows with the number of sampled paths, and notice the method needs answers that can be compared exactly.