Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

Chain of thought, as history

A prompt technique that turned out to be a capability

Wei et al. (arXiv:2201.11903, 2022) showed that giving a model a few worked examples in which the reasoning is written out — not just the answer — lifted accuracy sharply on multi-step arithmetic, commonsense and symbolic tasks. Kojima et al. (arXiv:2205.11916, 2022) then removed the examples entirely: appending the five words "Let's think step by step" to a zero-shot question produced most of the gain. Both results were surprising in the same way. Nothing was retrained; the only change was that the model emitted intermediate text before committing to an answer.

The mechanism is now uncontroversial. A decoder-only transformer performs a fixed amount of computation per emitted token, so a model that answers in one token has one token's worth of serial depth to work with. Writing the intermediate values out loud buys more serial computation, and it also changes what the later tokens attend to: the running result is in the context, so the next step conditions on it. Multi-step reliability is therefore multiplicative, which is the whole problem:

$$P(\text{chain correct}) \;=\; \prod_{i=1}^{N} p_i \;\approx\; p^{N}$$

At p = 0.66 per hop, a five-hop problem lands at 0.13. Raise the per-hop probability to 0.93 and the same five hops land at 0.70. Nothing about the model's knowledge changed — only whether it was allowed to externalise each intermediate step.

Toggle the trace and step the chain below. The bars are a seeded mock over 240 trials of the same arithmetic task; the chain is one representative trial, hop by hop.

Left: one trial's hops. Right: seeded accuracy with and without a written trace, over the same task.

💡 The durable idea: chain-of-thought is not a trick for extracting hidden knowledge. It is test-time compute — trading output tokens for serial depth and a working memory of intermediate results.
2

Self-consistency: sample many, take the majority

The first test-time-compute pattern

Wang et al. (arXiv:2203.11171, 2022) noticed that a single greedy chain is one sample from a distribution, and that its errors are not evenly distributed. Different chains fail in different places, so the correct final answer tends to be the one reached by the largest number of independent reasoning paths. Sample N chains at a non-zero temperature, take the majority vote over final answers, and accuracy climbs with N — no retraining, no reward model, just more sampling. This is the ancestor of every test-time-compute method that followed.

The arithmetic is a variance argument. If each chain is right with probability p > ½ and the chains are conditionally independent given the question, the majority of N chains is right with probability approaching 1 as N grows. The gains are front-loaded: most of the improvement arrives in the first handful of samples, then flattens, and the N chains cost N times the output tokens. Independence is the load-bearing assumption, and it is the first thing to break — a model that is systematically wrong about a question will produce N confidently wrong chains and the majority vote runs off the cliff with them.

Left: the tally of one sample of N chains. Right: majority-vote accuracy against N, from a seeded mock of 400 trials per point.

⚠️ The trap: majority voting assumes the errors are independent and that the majority is reachable. With a shared systematic misconception — a misread question, a poisoned fact in the context — more samples buy more confidence in the same wrong answer, and the cost is linear in N.
3

Spending thinking tokens

Accuracy against budget, and the point where more is worse

Self-consistency spends compute by sampling wider. A reasoning model spends it by thinking longer inside a single trajectory: the output is a long internal chain, and the model decides when to stop. Either way the same tradeoff appears. Accuracy rises with test-time compute and then saturates — a pattern often summarised as a log-linear curve, a + b·log(tokens), with a ceiling set by the task and the base capability. Past saturation, extra tokens buy latency and cost and nothing else.

There is a second, sharper effect that only shows up with trained reasoning models. A model trained with RL on verifiable rewards has learned how to allocate its thinking; a prompt that prescribes the shape of that thinking competes with the policy. Early in the budget the prescription can even help — it nudges a lazy trajectory — but as the budget grows it starts to override decisions the model was trained to make, and quality falls below the plain model at the same budget. The curve below makes that crossing explicit.

Seeded mock accuracy against thinking-token budget. The dashed marker is where a prescriptive prompt stops helping a reasoning model.

💡 The rule of thumb: spend more thinking tokens while the curve is still climbing, and stop when it flattens. The flattening point is a property of the task, not of the model, and it must be measured — a budget set by intuition is usually either wasteful or starved.
4

Prompting versus training

Where the prompt stops helping

The 2022 technique and the 2025 policy are not the same lever. A prompt reshapes the conditioning text of a fixed policy; a reasoning-trained model has internalised a search procedure into its weights through RLVR — reinforcement learning from verifiable rewards, with GRPO as the widely used optimiser. On easy tasks the two are hard to tell apart, because the prompt is enough. As difficulty rises, the prompt's ceiling appears: it can encourage more steps, but it cannot add a verification habit or a backtracking policy the model was never trained to have. That is why the same "think step by step" instruction is a large win on a base model and a coin flip on a reasoning model.

Note the caveat that the tuning source states plainly: RLVR trains on checkable outcomes, so a verifier that grades the final answer but not the reasoning invites reward hacking — the model learns to satisfy the checker rather than solve the task. A prompt cannot fix a policy trained against a weak checker.

Seeded mock accuracy against task difficulty, three ways of getting the answer: plain base model, base model plus a CoT prompt, and a reasoning-trained model.

5

When to think, and when not to

A decision, not a default

Reasoning is a cost centre with an uneven return, so the operational question is routing rather than maximising. Spend thinking tokens when the task is verifiable, multi-step, and wrong answers are expensive: arithmetic over retrieved figures, code that must pass tests, tool sequences whose intermediate state matters, or any problem where an intermediate result can be checked before it is used. Do not spend them on classification, extraction, formatting, rewriting or retrieval of a single fact — these have flat curves, and the extra tokens are pure latency.

The useful architecture is a cascade. Try the cheap path first, check it with something that is genuinely deterministic — a parser, a unit test, a schema validator, a calculator — and escalate to a larger budget only when the check fails. That is self-consistency generalised: the point was never "sample more", it was "have a way to tell which sample is right". Where no verifier exists, the honest position is that more thinking gives you a better-looking answer, not a more reliable one.

⚠️ Two failure modes in one line each: prescribing "think step by step" to a model trained to reason can degrade it by overriding its learned budget policy; and raising the thinking budget on an unverifiable task can buy confidence without correctness. Measure the curve before you spend on it.

Cheat sheet

QuestionThe answer that shapes the build
What does chain-of-thought actually buy?Serial computation and intermediate results in the context — test-time compute, not hidden knowledge.
Why does multi-step accuracy fall so fast?Reliability is multiplicative across hops: pN, so 0.9 per hop is 0.35 over ten hops.
What is self-consistency?Sample N chains, take the majority answer. The first test-time-compute pattern; accuracy rises with N then flattens.
When does majority voting fail?When errors are correlated — a shared misreading or a poisoned fact. More samples then amplify one wrong answer.
Why can "think step by step" hurt now?It competes with the budget policy an RL-trained model already learned, and it spends tokens on shape rather than search.
What did RLVR change?Verifiable rewards trained the search into the weights, so the model allocates its own thinking instead of following a prompt template.
Where should the thinking budget be spent?Verifiable, multi-step, high-cost-of-error tasks. Route cheap tasks to cheap paths and escalate on a failed check.

Further reading

6

Check your understanding

0/5 answered