Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

From logits to a distribution

Softmax, then a temperature dial

At each step the model produces one raw score, a logit, for every token in its vocabulary. Logits are unbounded and unnormalised; softmax turns them into a probability distribution by exponentiating each score and dividing by the sum. Temperature divides the logits before the exponential. A temperature below 1 sharpens the distribution — differences between scores are magnified, so the leader pulls away. A temperature above 1 flattens it, giving unlikely tokens more of a chance. Temperature does not change which token the model prefers; it changes how strongly that preference dominates.

Below is one fixed logit vector over eight candidate next words. Move the temperature and watch the bars; the other two sliders, top-k and top-p, are the truncation rules the next section takes up. The button draws one sample from the surviving distribution — through a seeded stream, so a reload reproduces the same draw.

Probability per candidate after softmax and truncation. Grey bars were cut; the labelled bar is the sampled token.

💡 Durable idea: the model gives you a ranking plus a shape; the sampler is yours to choose. Every "creative" versus "factual" setting is a decision about how much of that shape you keep.
2

Top-k, top-p, and greedy

Truncating the tail reshapes what is left

Two families of rule trim the distribution before sampling. Top-k keeps the k highest-probability tokens and discards the rest, then renormalises. It is simple, but the fixed k ignores shape: when the model is certain it unnecessarily admits the runners-up, and when it is genuinely uncertain it truncates a tail that might have contained the right answer. Top-p, or nucleus sampling, keeps the smallest set of tokens whose cumulative probability reaches p and drops everything beyond it. The set grows when the model is unsure and shrinks when it is confident, which is why it is the more common default.

Truncation and temperature interact. Lowering the temperature concentrates mass on the leader, which can pull the nucleus tight on its own; raising the top-p ceiling lets more of the long tail back in. The two are applied together in the demo above — a token survives only if it is inside both the top-k prefix and the nucleus, and the survivors are renormalised to sum to one.

Greedy / temperature 0
Always take the argmax

Deterministic in exact arithmetic, repetitive in practice. It is what "temperature 0" is supposed to mean.

Top-p, then sample
Keep the nucleus, sample within it

The tail is removed but the head is sampled in proportion to probability, so the output varies.

⚠️ Trap: sampling parameters are not a "quality dial". A high temperature cannot make a weak model correct, and a low one cannot make it diverse. They trade variance for fidelity within whatever distribution the model already has.
3

Temperature 0 is not deterministic

Same prompt, same setting, different output

Temperature 0 removes the sampler's randomness. It does not remove the rest of the machine's. Two calls with the same prompt can still disagree because the arithmetic that produces the logits is not the same arithmetic each time. Modern serving stacks batch many requests together, and the size and composition of that batch changes the order in which floating-point reductions happen; floating-point addition is not associative, so a different order is a different number. Mixture-of-experts routing depends on the batch as well, sending the same token to a slightly different set of experts. A provider can also change weights, kernels or a default between your two calls without telling you, and a version bump is a different function.

The grid below is eight calls of one prompt at temperature 0, shown as their emitted token ids. Most positions agree — that is the greedy argmax doing its job — and a handful drift, which is the infrastructure underneath. This is why a temperature-0 test that passes today can fail tomorrow, and why evaluation harnesses pin a model version and a serving configuration rather than trusting the setting.

Eight calls, ten positions each. Highlighted cells are positions where the batch produced a different token.

The drift is not sampled here; it is drawn from a seeded stream to stand in for the sources listed under the grid. Nothing in a temperature-0 call is guaranteed bit-reproducible across batches, hardware, or provider versions.

💡 Practical rule: if you need reproducibility, control the whole path — pinned model version, pinned serving configuration, batch-independent kernels — or accept that "deterministic" was always an aspiration, not a property.
4

Sampling the same prompt many times

Variance is information

If you sample the same question many times, the number of distinct answers is a rough proxy for how much the model actually knows. At temperature 0 every draw returns the same answer, which tells you nothing about whether that answer is stable. As temperature rises, the sampler explores more of the distribution; a question with a crisp, well-supported answer keeps returning one or two answers, while a question the model is guessing on scatters across many. This is the seed of self-consistency: sample several reasoning paths, keep the majority answer, and use the disagreement as a signal.

The chart below samples repeatedly from a small five-answer distribution and counts how many distinct answers appear as the temperature rises. It is not a real model — the point is the shape — but the shape is the reason self-consistency works on tasks with a checkable answer.

Distinct answers (out of five) versus temperature, over the number of draws you choose.

⚠️ Trap: agreement is not correctness. Five samples converging on the same wrong answer is a confident mistake, and self-consistency will happily preserve it. Judge the majority against evidence, not against itself.
5

Logprobs and the calibration problem

A confidence signal you should not over-trust

Many APIs will return logprobs — the log-probability the model assigned to each token it emitted, and often the top alternatives. That is a genuine internal signal: a token with logprob near zero was almost forced, and a token with a much lower logprob was one of several live options. It is genuinely useful for flagging uncertain generations, for constrained decoding, and for detecting when the model was guessing between plausible continuations.

The trap is treating that number as a probability that the answer is correct. A reliability diagram plots stated confidence against observed accuracy; a perfectly calibrated model lies on the diagonal. Post-training on human preferences tends to push models off it, and the direction is usually optimism: the model says 0.9 and is right far less often. The curve below is a seeded mock of that overconfidence, not a measurement of any particular model, but the phenomenon is well documented — reinforcement learning from human feedback optimises for answers a rater prefers, and a confident tone is often preferred.

Stated confidence (x) versus empirical accuracy (y). The dotted diagonal is perfect calibration.

💡 Carry this forward: logprobs are a useful internal uncertainty signal and a poor external correctness signal. Calibrate them against a labelled set, or use them only to decide when to ask a human, retrieve more, or abstain.

Cheat sheet

QuestionShort answer
What turns logits into probabilities?Softmax: exponentiate and normalise. Temperature divides the logits first.
What does temperature below 1 do?Sharpens the distribution; the leader's advantage is magnified.
Top-k versus top-p?Top-k keeps a fixed count; top-p keeps the smallest set reaching cumulative mass p. Top-p adapts to confidence.
Does temperature 0 guarantee the same output?No. Batched floating-point reduction order, MoE routing, hardware and provider version changes all break the tie.
What are logprobs good for?Flagging uncertain generations and comparing alternatives. Not for reading off the probability of being correct.
How do post-RLHF models calibrate?Often badly and optimistically — stated confidence runs ahead of accuracy.
Why sample many times?Distinct-answer count is a cheap uncertainty signal, and a majority vote can beat a single greedy pass.

Further reading

6

Check your understanding

0/4 answered