Sampling, nondeterminism and confidence
A language model does not emit a word. It emits a vector of scores over its whole vocabulary, and something downstream picks one. That picking step — the sampler — decides how creative the output looks, whether it is reproducible, and how much the model's own probabilities can be trusted. This part takes the sampler apart: softmax and temperature, top-k and top-p, the myth of the deterministic setting, and the uncomfortable gap between a model's stated confidence and its accuracy.
From logits to a distribution
Softmax, then a temperature dial
At each step the model produces one raw score, a logit, for every token in its vocabulary. Logits are unbounded and unnormalised; softmax turns them into a probability distribution by exponentiating each score and dividing by the sum. Temperature divides the logits before the exponential. A temperature below 1 sharpens the distribution — differences between scores are magnified, so the leader pulls away. A temperature above 1 flattens it, giving unlikely tokens more of a chance. Temperature does not change which token the model prefers; it changes how strongly that preference dominates.
Below is one fixed logit vector over eight candidate next words. Move the temperature and watch the bars; the other two sliders, top-k and top-p, are the truncation rules the next section takes up. The button draws one sample from the surviving distribution — through a seeded stream, so a reload reproduces the same draw.
Probability per candidate after softmax and truncation. Grey bars were cut; the labelled bar is the sampled token.
Top-k, top-p, and greedy
Truncating the tail reshapes what is left
Two families of rule trim the distribution before sampling. Top-k keeps the k highest-probability tokens and discards the rest, then renormalises. It is simple, but the fixed k ignores shape: when the model is certain it unnecessarily admits the runners-up, and when it is genuinely uncertain it truncates a tail that might have contained the right answer. Top-p, or nucleus sampling, keeps the smallest set of tokens whose cumulative probability reaches p and drops everything beyond it. The set grows when the model is unsure and shrinks when it is confident, which is why it is the more common default.
Truncation and temperature interact. Lowering the temperature concentrates mass on the leader, which can pull the nucleus tight on its own; raising the top-p ceiling lets more of the long tail back in. The two are applied together in the demo above — a token survives only if it is inside both the top-k prefix and the nucleus, and the survivors are renormalised to sum to one.
Deterministic in exact arithmetic, repetitive in practice. It is what "temperature 0" is supposed to mean.
The tail is removed but the head is sampled in proportion to probability, so the output varies.
Temperature 0 is not deterministic
Same prompt, same setting, different output
Temperature 0 removes the sampler's randomness. It does not remove the rest of the machine's. Two calls with the same prompt can still disagree because the arithmetic that produces the logits is not the same arithmetic each time. Modern serving stacks batch many requests together, and the size and composition of that batch changes the order in which floating-point reductions happen; floating-point addition is not associative, so a different order is a different number. Mixture-of-experts routing depends on the batch as well, sending the same token to a slightly different set of experts. A provider can also change weights, kernels or a default between your two calls without telling you, and a version bump is a different function.
The grid below is eight calls of one prompt at temperature 0, shown as their emitted token ids. Most positions agree — that is the greedy argmax doing its job — and a handful drift, which is the infrastructure underneath. This is why a temperature-0 test that passes today can fail tomorrow, and why evaluation harnesses pin a model version and a serving configuration rather than trusting the setting.
Eight calls, ten positions each. Highlighted cells are positions where the batch produced a different token.
The drift is not sampled here; it is drawn from a seeded stream to stand in for the sources listed under the grid. Nothing in a temperature-0 call is guaranteed bit-reproducible across batches, hardware, or provider versions.
Sampling the same prompt many times
Variance is information
If you sample the same question many times, the number of distinct answers is a rough proxy for how much the model actually knows. At temperature 0 every draw returns the same answer, which tells you nothing about whether that answer is stable. As temperature rises, the sampler explores more of the distribution; a question with a crisp, well-supported answer keeps returning one or two answers, while a question the model is guessing on scatters across many. This is the seed of self-consistency: sample several reasoning paths, keep the majority answer, and use the disagreement as a signal.
The chart below samples repeatedly from a small five-answer distribution and counts how many distinct answers appear as the temperature rises. It is not a real model — the point is the shape — but the shape is the reason self-consistency works on tasks with a checkable answer.
Distinct answers (out of five) versus temperature, over the number of draws you choose.
Logprobs and the calibration problem
A confidence signal you should not over-trust
Many APIs will return logprobs — the log-probability the model assigned to each token it emitted, and often the top alternatives. That is a genuine internal signal: a token with logprob near zero was almost forced, and a token with a much lower logprob was one of several live options. It is genuinely useful for flagging uncertain generations, for constrained decoding, and for detecting when the model was guessing between plausible continuations.
The trap is treating that number as a probability that the answer is correct. A reliability diagram plots stated confidence against observed accuracy; a perfectly calibrated model lies on the diagonal. Post-training on human preferences tends to push models off it, and the direction is usually optimism: the model says 0.9 and is right far less often. The curve below is a seeded mock of that overconfidence, not a measurement of any particular model, but the phenomenon is well documented — reinforcement learning from human feedback optimises for answers a rater prefers, and a confident tone is often preferred.
Stated confidence (x) versus empirical accuracy (y). The dotted diagonal is perfect calibration.
Cheat sheet
| Question | Short answer |
|---|---|
| What turns logits into probabilities? | Softmax: exponentiate and normalise. Temperature divides the logits first. |
| What does temperature below 1 do? | Sharpens the distribution; the leader's advantage is magnified. |
| Top-k versus top-p? | Top-k keeps a fixed count; top-p keeps the smallest set reaching cumulative mass p. Top-p adapts to confidence. |
| Does temperature 0 guarantee the same output? | No. Batched floating-point reduction order, MoE routing, hardware and provider version changes all break the tie. |
| What are logprobs good for? | Flagging uncertain generations and comparing alternatives. Not for reading off the probability of being correct. |
| How do post-RLHF models calibrate? | Often badly and optimistically — stated confidence runs ahead of accuracy. |
| Why sample many times? | Distinct-answer count is a cheap uncertainty signal, and a majority vote can beat a single greedy pass. |
Further reading
- Fan, Lewis & Dauphin, "Hierarchical Neural Story Generation", ACL 2018 (arXiv:1805.04833) — the paper that introduced top-k sampling for open-ended generation.
- Holtzman et al., "The Curious Case of Neural Text Degeneration", ICLR 2020 (arXiv:1904.09751) — nucleus sampling (top-p) and the argument against pure truncation by rank.
- Guo et al., "On Calibration of Modern Neural Networks", ICML 2017 (arXiv:1706.04599) — reliability diagrams and expected calibration error, the tools this page borrows.
- Kadavath et al., "Language Models (Mostly) Know What They Know", 2022 (arXiv:2207.05221) — models carry a usable internal sense of their own correctness, but it must be elicited and calibrated.
- OpenAI, "GPT-4 Technical Report", 2023 (arXiv:2303.08774) — the observation that post-training on preferences reduces calibration in the direction of overconfidence.
- Tian et al., "Just Ask for Calibration", EMNLP 2023 (arXiv:2305.14975) — eliciting calibrated probabilities in words rather than logprobs.
- Thinking Machines, "Defeating Nondeterminism in LLM Inference", 2025 — a concrete account of why greedy decoding drifts at the level of the batch and the kernel.