Skip to content
Road to Intelligence

Concept · Chapter 14: Reasoning Models

Self-Consistency

Must knowKnow well10 minDifficulty

Self-consistency samples several independent reasoning paths for the same question and returns the final answer they most often agree on.

The problem

One sampled or greedy chain of thought can go wrong at any step, and its single answer carries no sign of how reliable it is.

The solution

Sample many chains at a non-zero temperature, extract each final answer, and take a plurality vote; disagreement doubles as a signal of uncertainty.

The consequence

Accuracy rises with the number of samples on tasks with checkable final answers, at a proportional inference cost, but a vote amplifies whatever mistake the samples share.

Vote on answers, not on wording

Two chains can be worded completely differently and still end in 21. Self-consistency ignores the wording and counts final answers. If eight samples end in 21, 42, 21, 20, 21, 42, 21 and 21, the plurality answer is 21 with five votes.

Wang and colleagues proposed replacing greedy chain-of-thought decoding with sampling plus answer aggregation, reporting gains including +17.9% on GSM8K, +11.0% on SVAMP and +12.2% on AQuA. Established No training is involved; it is purely an inference-time procedure.

The intuition: a hard problem usually has several valid routes to one correct answer, while mistakes scatter across many different wrong answers. Correct paths agree with each other; wrong paths mostly disagree.

Availability is not selection

If each independent sample is correct with probability pp, the chance that at least one of NN is correct is

P(at least one correct)=1−(1−p)N.P(\text{at least one correct}) = 1 - (1 - p)^N.

Tiny example. With p=0.25p = 0.25 and N=4N = 4: 1−0.754=1−0.316=0.6841 - 0.75^4 = 1 - 0.316 = 0.684. A correct answer is present 68% of the time. But the vote returns the most common answer, and with p=0.25p = 0.25 a single dominant wrong answer can easily be more common than the right one. Availability measures what a perfect selector could do; the vote is a specific, imperfect selector.

The same formula fails in a second way: it assumes every attempt is an independent coin flip with the same pp. Plugging in an average accuracy over a mix of easy and hard tasks overstates the gain, because on the hard tasks pp is near zero however many samples you draw.

When agreement misleads

In the consensus lab, the commonest mistake (forgetting to halve, giving 42) can outvote the correct 21 even when 21 appears several times. Push “Copy the first answer” to 100% and every candidate repeats the first one: thirty-two votes, one opinion.

Real models need no literal copying to fail together. Samples share a model, a prompt and a training history, so they can share a misreading of the question. A unanimous vote on a misread question is confident and wrong. The fix is diversity of evidence, not more samples from the same source: different prompts, tools, or an independent check.

Practical notes

  • Cost. NN samples cost about NN times the generation tokens. Samples can run in parallel, so latency grows less than total compute does.
  • Ties. Decide in advance. The lab abstains on a tie rather than picking arbitrarily.
  • Answer extraction. Normalise answers before counting (“21”, “21.0” and “twenty-one” are one answer).
  • Confidence. The vote share is a useful uncertainty signal: 15 of 16 agreeing is different from 6 of 16. It is not a calibrated probability.
  • Free-form outputs such as essays or code have no exact answer to count; they need a scorer or tests, see verifiers.

OpenAI reported o1 at 74% on AIME 2024 with one sample per problem and 83% with a vote over 64 samples. Established Those are the company's own figures, and they illustrate the trade: about 64 times the samples for nine more points.

Mini experiment

In the lab, keep the default settings and press New seed ten times, playing each batch. Tally how often the vote returns 21, 42 or a tie, and how often the verifier's pick is 21. Then set the candidate budget to 32 and repeat. Which selector benefits more from more candidates at this accuracy?

What to remember

  • Sample N chains, extract final answers, return the most common one.
  • Wang et al. reported +17.9% on GSM8K over greedy chain-of-thought decoding.
  • 1 − (1 − p)^N is the chance some sample is right, not the chance the vote is right.
  • Shared mistakes win votes: correlated samples add cost without adding evidence.
  • Voting needs answers that can be compared exactly; free-form text needs another selector.

Key papers

Essential

Self-Consistency Improves Chain of Thought Reasoning in Language Models

Xuezhi Wang, Jason Wei et al. · 2022 · ICLR 2023

Turned extra inference compute into accuracy without any training: sample many reasoning paths and keep the answer they most often reach.

How to read it: Look at how accuracy grows with the number of sampled paths, and notice the method needs answers that can be compared exactly.

~35 min readarXiv:2203.11171✓ verified 2026-10-06