Reasoning: CoT, self-consistency, and what reasoning models changed
For three years the highest-leverage sentence in a prompt was "let's think step by step". It still works on models that were never trained to reason, and it is now frequently the wrong thing to say to a model that was. This part treats chain-of-thought as history rather than advice: what the 2022 papers actually showed, how self-consistency turned sampling into the first test-time-compute pattern, and what changed when labs started training the search into the weights with reinforcement learning on verifiable rewards.
Chain of thought, as history
A prompt technique that turned out to be a capability
Wei et al. (arXiv:2201.11903, 2022) showed that giving a model a few worked examples in which the reasoning is written out — not just the answer — lifted accuracy sharply on multi-step arithmetic, commonsense and symbolic tasks. Kojima et al. (arXiv:2205.11916, 2022) then removed the examples entirely: appending the five words "Let's think step by step" to a zero-shot question produced most of the gain. Both results were surprising in the same way. Nothing was retrained; the only change was that the model emitted intermediate text before committing to an answer.
The mechanism is now uncontroversial. A decoder-only transformer performs a fixed amount of computation per emitted token, so a model that answers in one token has one token's worth of serial depth to work with. Writing the intermediate values out loud buys more serial computation, and it also changes what the later tokens attend to: the running result is in the context, so the next step conditions on it. Multi-step reliability is therefore multiplicative, which is the whole problem:
At p = 0.66 per hop, a five-hop problem lands at 0.13. Raise the per-hop probability to 0.93 and the same five hops land at 0.70. Nothing about the model's knowledge changed — only whether it was allowed to externalise each intermediate step.
Toggle the trace and step the chain below. The bars are a seeded mock over 240 trials of the same arithmetic task; the chain is one representative trial, hop by hop.
Left: one trial's hops. Right: seeded accuracy with and without a written trace, over the same task.
Self-consistency: sample many, take the majority
The first test-time-compute pattern
Wang et al. (arXiv:2203.11171, 2022) noticed that a single greedy chain is one sample from a distribution, and that its errors are not evenly distributed. Different chains fail in different places, so the correct final answer tends to be the one reached by the largest number of independent reasoning paths. Sample N chains at a non-zero temperature, take the majority vote over final answers, and accuracy climbs with N — no retraining, no reward model, just more sampling. This is the ancestor of every test-time-compute method that followed.
The arithmetic is a variance argument. If each chain is right with probability p > ½ and the chains are conditionally independent given the question, the majority of N chains is right with probability approaching 1 as N grows. The gains are front-loaded: most of the improvement arrives in the first handful of samples, then flattens, and the N chains cost N times the output tokens. Independence is the load-bearing assumption, and it is the first thing to break — a model that is systematically wrong about a question will produce N confidently wrong chains and the majority vote runs off the cliff with them.
Left: the tally of one sample of N chains. Right: majority-vote accuracy against N, from a seeded mock of 400 trials per point.
Spending thinking tokens
Accuracy against budget, and the point where more is worse
Self-consistency spends compute by sampling wider. A reasoning model spends it by thinking longer inside a single trajectory: the output is a long internal chain, and the model decides when to stop. Either way the same tradeoff appears. Accuracy rises with test-time compute and then saturates — a pattern often summarised as a log-linear curve, a + b·log(tokens), with a ceiling set by the task and the base capability. Past saturation, extra tokens buy latency and cost and nothing else.
There is a second, sharper effect that only shows up with trained reasoning models. A model trained with RL on verifiable rewards has learned how to allocate its thinking; a prompt that prescribes the shape of that thinking competes with the policy. Early in the budget the prescription can even help — it nudges a lazy trajectory — but as the budget grows it starts to override decisions the model was trained to make, and quality falls below the plain model at the same budget. The curve below makes that crossing explicit.
Seeded mock accuracy against thinking-token budget. The dashed marker is where a prescriptive prompt stops helping a reasoning model.
Prompting versus training
Where the prompt stops helping
The 2022 technique and the 2025 policy are not the same lever. A prompt reshapes the conditioning text of a fixed policy; a reasoning-trained model has internalised a search procedure into its weights through RLVR — reinforcement learning from verifiable rewards, with GRPO as the widely used optimiser. On easy tasks the two are hard to tell apart, because the prompt is enough. As difficulty rises, the prompt's ceiling appears: it can encourage more steps, but it cannot add a verification habit or a backtracking policy the model was never trained to have. That is why the same "think step by step" instruction is a large win on a base model and a coin flip on a reasoning model.
Note the caveat that the tuning source states plainly: RLVR trains on checkable outcomes, so a verifier that grades the final answer but not the reasoning invites reward hacking — the model learns to satisfy the checker rather than solve the task. A prompt cannot fix a policy trained against a weak checker.
Seeded mock accuracy against task difficulty, three ways of getting the answer: plain base model, base model plus a CoT prompt, and a reasoning-trained model.
When to think, and when not to
A decision, not a default
Reasoning is a cost centre with an uneven return, so the operational question is routing rather than maximising. Spend thinking tokens when the task is verifiable, multi-step, and wrong answers are expensive: arithmetic over retrieved figures, code that must pass tests, tool sequences whose intermediate state matters, or any problem where an intermediate result can be checked before it is used. Do not spend them on classification, extraction, formatting, rewriting or retrieval of a single fact — these have flat curves, and the extra tokens are pure latency.
The useful architecture is a cascade. Try the cheap path first, check it with something that is genuinely deterministic — a parser, a unit test, a schema validator, a calculator — and escalate to a larger budget only when the check fails. That is self-consistency generalised: the point was never "sample more", it was "have a way to tell which sample is right". Where no verifier exists, the honest position is that more thinking gives you a better-looking answer, not a more reliable one.
Cheat sheet
| Question | The answer that shapes the build |
|---|---|
| What does chain-of-thought actually buy? | Serial computation and intermediate results in the context — test-time compute, not hidden knowledge. |
| Why does multi-step accuracy fall so fast? | Reliability is multiplicative across hops: pN, so 0.9 per hop is 0.35 over ten hops. |
| What is self-consistency? | Sample N chains, take the majority answer. The first test-time-compute pattern; accuracy rises with N then flattens. |
| When does majority voting fail? | When errors are correlated — a shared misreading or a poisoned fact. More samples then amplify one wrong answer. |
| Why can "think step by step" hurt now? | It competes with the budget policy an RL-trained model already learned, and it spends tokens on shape rather than search. |
| What did RLVR change? | Verifiable rewards trained the search into the weights, so the model allocates its own thinking instead of following a prompt template. |
| Where should the thinking budget be spent? | Verifiable, multi-step, high-cost-of-error tasks. Route cheap tasks to cheap paths and escalate on a failed check. |
Further reading
- Wei et al., "Chain-of-Thought Prompting Elicits Reasoning in Large Language Models", NeurIPS 2022 — the few-shot worked-example result.
- Kojima et al., "Large Language Models are Zero-Shot Reasoners", NeurIPS 2022 — "Let's think step by step", with no examples at all.
- Wang et al., "Self-Consistency Improves Chain of Thought Reasoning in Language Models", ICLR 2023 — the majority-vote-over-samples pattern.
- Lightman et al., "Let's Verify Step by Step", 2023 — process supervision: grading the reasoning, not only the answer.
- Snell et al., "Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters", 2024 — when spending at inference beats a bigger base model.
- DeepSeek-AI, "DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning", 2025 — RLVR and GRPO at scale, including the emergent long chain.
- Anthropic, "Building effective agents", December 2024 — the evaluation-and-check pattern the cascade builds on.