Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

The question

Where does the probability actually live?

Ask a working engineer where probability sits inside a neural network and the answer usually points at the loss. That is half right. The loss is where probability is used, but the probability is produced earlier, by an output layer that turns a vector of real numbers into something that sums to one. For a classification head that layer is the softmax; for a language model it is the same softmax applied to a vocabulary-sized vector of logits; for a Gaussian head it is a mean and a variance. Every one of those is a distribution parameterised by the network, and the network's job is to move that distribution onto the data.

Once you see the output as a distribution, the training objective stops being arbitrary. The likelihood of the observed label under the predicted distribution is a number between zero and one; training maximises it. Maximising a product is the same as minimising the sum of negative logs, which is why the loss is written with a minus sign and a logarithm. Cross-entropy is not a separate invention bolted onto the softmax — it is the negative log-likelihood of the categorical model the softmax defines. Flip the question around and the same identity explains why the loss is large when the model is confidently wrong: the predicted probability of the true label was tiny, and the negative log of a tiny number is large.

Two further questions follow from that picture, and they are where the probabilistic reading earns its keep. First, is the number the model reports true? A network that says 0.9 should be right about ninety percent of the time on that kind of input, and if it is not, the distribution is miscalibrated even when the accuracy is fine. Calibration is a statement about the distribution, not the decision, and it has its own diagnostic — the reliability diagram — and its own cheap fix, temperature scaling. Second, how much should a single prediction be believed at all? A point estimate throws away the shape of the distribution. Gaussian processes, variational autoencoders and diffusion models each preserve it in a different way: a closed-form posterior, a learned latent code, or a learned reverse process.

This part is the volume's cashing-in chapter. The softmax is a distribution built in the information theory part; the likelihood is Part 1; the latent-variable ELBO is Part 12; the Gaussian conditioning behind the GP is the multivariate Gaussian part of Volume I. Nothing new is introduced except the assembly. The plan is to take the output distribution apart, then the loss, then the calibration of the loss, then uncertainty, and finally the generative models that make the whole distribution the object being learned.

💡 By the end of this part you'll see why the softmax is a categorical distribution with a temperature, why cross-entropy is exactly the negative log-likelihood of the true label, how reliability diagrams expose overconfidence and temperature scaling repairs it, how a Gaussian process shrinks its uncertainty near data, and how VAEs and diffusion models learn transformations of distributions rather than single outputs.
2

Softmax as a distribution

A dial that controls how committed the model is

A network's final linear layer hands you a vector $z$ of real numbers called logits. They have no units and can be any size or sign; the only thing that matters is their differences, because adding a constant to every logit changes nothing about the ranking. To turn them into probabilities you exponentiate and normalise. The exponential makes every entry positive and makes differences multiplicative, and the sum in the denominator forces the result to add to one. That is the whole definition:

$$\operatorname{softmax}(z)_i=\frac{e^{z_i}}{\sum_{j=1}^{K}e^{z_j}},\qquad \operatorname{softmax}(z/T)_i=\frac{e^{z_i/T}}{\sum_{j=1}^{K}e^{z_j/T}}.$$

The second expression introduces a temperature $T$. Dividing every logit by $T$ before the softmax reshapes the distribution without changing which class wins. When $T$ is small the logits are stretched apart, the largest entry dominates, and the distribution becomes nearly one-hot — the model is confident. When $T$ is large the logits are squashed together and the distribution flattens toward the uniform $1/K$, which has the maximum entropy any $K$-class distribution can have. This one dial is the difference between a greedy sampler and a creative one.

You have already met the dial in the serving guide. The decode loop samples each token from the softmax over the next-token logits, and the temperature is the knob that trades diversity for reliability. Set it to zero and decoding becomes greedy argmax; raise it and the tail of the distribution gets sampled. Speculative decoding takes the same distribution and uses a small draft model to propose tokens that a larger model then accepts or rejects, and the accept-reject test is written in terms of these probabilities — it is a couple of distributions being compared, nothing more.

Drag the four logits and the temperature. Two things are worth watching. The ranking of the bars is fixed by the logits; the temperature only changes how spread out they are. And as $T$ falls the entropy drops toward zero, while as $T$ rises it climbs toward the uniform ceiling $\log K$. The distribution is doing one smooth thing and the readout names it.

Bars are the class probabilities $\operatorname{softmax}(z/T)$. The dashed line is the uniform distribution $1/K$, the flattest this softmax can get.

The distribution the softmax defines is called categorical, and its entropy is the number of nats of information the model is undecided by. That is not a metaphor: a distribution with entropy $H$ costs at least $H$ nats to encode, so an uncertain model is literally an expensive one. The next section makes that accounting exact by scoring the distribution against what actually happened.

3

Cross-entropy is negative log-likelihood

The loss is just the log of the probability you assigned to the truth

Suppose the true class is $y$ and the model predicted the distribution $q$. The likelihood of the observation under the model is $q(y)$, a single number. Training wants that number to be as large as possible, and because logs turn products into sums, the equivalent objective is to minimise its negative logarithm. One example, one term:

$$\mathrm{NLL}(y,q)=-\log q(y),\qquad H(p,q)=-\sum_{i=1}^{K}p_i\log q_i,\qquad H(p,q)=H(p)+\mathrm{KL}(p\,\|\,q).$$

The middle expression is the cross-entropy between a true distribution $p$ and the model's $q$. When the label is a one-hot vector — all its mass on the true class — the sum collapses to the single term $-\log q(y)$, which is the negative log-likelihood. The identity on the right is the reason cross-entropy is the standard loss: it equals the entropy of the truth, which the model cannot change, plus the KL divergence from truth to model. Minimising cross-entropy therefore minimises the KL divergence, the same quantity the information theory part read as an excess codelength. And by Gibbs' inequality the KL is never negative, so the loss bottoms out at the true entropy and no lower.

The shape of the function carries the practical lesson. Because the logarithm is steep near zero, a confident mistake is punished hard: predicting $q(y)=0.01$ costs about $4.6$ nats, while $q(y)=0.5$ costs only $0.69$. A model that puts almost no mass on the truth is driven by an enormous gradient, which is exactly the pressure that makes a classifier stop being confidently wrong. It also explains why numerical implementations clamp the log or use a fused log-softmax: the raw logarithm of a probability that underflowed to zero is not a number you want to compute.

In the demo the model's predicted distribution is fixed to one shape — a large logit on class 0, zeros elsewhere — and you choose which class is the truth. When the truth is class 0 the loss is small and falls as the margin grows. When the truth is any other class the model is confidently wrong, $q(y)$ is minuscule, and the loss explodes. Move the margin slider with the truth set away from the peak and watch the meter fill.

Bars are the predicted distribution $q$. The dashed line marks the true class $y$; the loss is $-\log q(y)$, read from the height of the bar the line sits on.

This is the same object the language-model guide calls the training loss and the evaluation guide calls perplexity. Perplexity is just $e^{\mathrm{NLL}}$, the effective number of equally likely choices the model was entertaining; a perplexity of $K$ means the model was as confused as a uniform distribution over $K$ tokens. And preference training in the RLHF part starts from the log probabilities the same softmax produced, so the whole pipeline from pretraining to alignment is arithmetic on a categorical distribution.

4

Calibration and reliability

A probability is a promise about frequency

A classifier that outputs $0.9$ is making a falsifiable claim: on all the inputs it labels with confidence near $0.9$, about ninety percent should be correct. A model that satisfies this is calibrated. Calibration is separate from accuracy. A model can be accurate and wildly overconfident (right on most inputs, but claiming $0.99$ when it is right only $0.8$ of the time), or accurate and underconfident. Modern deep networks are usually the first: they win on accuracy and then report confidences no one should trust.

The diagnostic is a reliability diagram. Bin the predictions by their stated confidence, then for each bin compute the empirical accuracy — the fraction of those predictions that were actually correct. Plot accuracy against confidence. A perfectly calibrated model lies on the diagonal $y=x$, because its claimed confidence equals the observed frequency. Points below the diagonal are overconfident; points above are underconfident. The gap between the curve and the diagonal, weighted by how many predictions fall in each bin, is the expected calibration error:

$$\mathrm{ECE}=\sum_{b=1}^{B}\frac{n_b}{N}\,\bigl|\,\mathrm{acc}_b-\mathrm{conf}_b\,\bigr|,\qquad \mathrm{conf}_b=\frac{1}{n_b}\sum_{i\in b}\hat p_i,\qquad \mathrm{acc}_b=\frac{1}{n_b}\sum_{i\in b}\mathbf{1}\{\hat y_i=y_i\}.$$

Overconfidence has a remarkably cheap repair called temperature scaling. Leave the trained network alone and divide its logits by a single number $T>1$ before the softmax. This softens every distribution, pulling the confidence down toward the accuracy without changing a single predicted class, because dividing by a positive number does not reorder the logits. You fit the one parameter $T$ on a held-out set by minimising the same negative log-likelihood, and the reliability curve walks back to the diagonal. One scalar repairs a network with millions of weights.

The demo simulates a classifier whose predictions are warped by an overconfidence factor. The true probability of the positive class is drawn first, the label is sampled from it, and the model's confidence is a sharpened version of the truth, so its errors are exactly the overconfidence you would see in practice. The pink points are the raw model; the blue points are the same model after temperature scaling. Slide the temperature and watch the blue curve snap onto the diagonal while the class decisions never move.

Reliability diagram: predicted confidence on the $x$-axis, empirical accuracy on the $y$-axis. Pink is the raw model, blue is temperature-scaled; the pink verticals are the calibration gaps that add up to the ECE.

Calibration is not the same as a good decision rule. If all you need is the argmax, a miscalibrated model may serve perfectly well; if the probability is going to be multiplied by a cost, fed into a Bayesian filter, or used to decide whether to hand a task to a human, then the number itself carries weight and ECE is the thing to watch. This is also why evaluation sets double as calibration sets in the evaluation part: the same held-out data that measures accuracy measures the honesty of the confidences.

5

Uncertainty: GPs, VAEs, diffusion

Keeping the whole distribution instead of one point

A point prediction throws the distribution away at the last moment. Three families of model refuse to. The first is the Gaussian process, which puts a prior over functions: any finite set of function values is jointly Gaussian, and conditioning on observations gives a posterior that is itself Gaussian everywhere. With an RBF kernel and a handful of noisy observations, the posterior mean is a smooth interpolation and the posterior variance has a shape you can read: it collapses toward the noise floor at each observed input and swells back toward the prior far away from all of them. The formulas are the multivariate Gaussian conditioning of Volume I, written for an infinite index set:

$$\mu_*(x)=k(x)^\top\!\left(K+\sigma_n^2 I\right)^{-1}\!y,\qquad \sigma_*^2(x)=k(x,x)-k(x)^\top\!\left(K+\sigma_n^2 I\right)^{-1}\!k(x),\qquad k(x,x')=\sigma_f^2 e^{-(x-x')^2/2\ell^2}.$$

Drag the lengthscale and the noise and watch the band. A long lengthscale makes the function smooth and the uncertainty broad; a short one lets the posterior wiggle and the band pinch tightly at every observation. More noise raises the floor and keeps the band from ever reaching zero, because the observations themselves are not trusted. The important thing is not the curve — it is the second line of the plot, the variance, which no point-estimate model can draw at all.

The line is the posterior mean; the shaded band is $\pm 2\sigma_*$. Markers are the noisy observations. The band pinches at each marker and widens back toward the prior between them.

The second family is the variational autoencoder, which buys a distribution over latent codes by learning the transformation itself. An encoder maps a data point $x$ to a Gaussian $q(z\mid x)$ in a low-dimensional latent space, a decoder maps a code $z$ back to a distribution over data, and training maximises the evidence lower bound from Part 12: a reconstruction term plus a KL term that keeps the codes near a standard normal. Sampling from that latent Gaussian and decoding gives a whole cloud of plausible outputs, and the KL term is what makes the latent space smooth enough for interpolation to mean something.

The third family is diffusion, which learns a distribution by learning to destroy and then rebuild structure. The forward process is a fixed chain of Gaussians that adds a little noise at each step, and the reverse process is a network trained to undo one step at a time. The forward step is the simple object, and once you see it the rest of the pipeline is a matter of parameterisation:

$$q(x_t\mid x_{t-1})=\mathcal{N}\!\left(x_t;\sqrt{1-\beta_t}\,x_{t-1},\;\beta_t I\right),\qquad q(x_t\mid x_0)=\mathcal{N}\!\left(x_t;\sqrt{\bar\alpha_t}\,x_0,\;(1-\bar\alpha_t)I\right).$$

Each step keeps the signal coefficient $\sqrt{1-\beta_t}$ and adds variance $\beta_t$, so the chain is variance-preserving and the marginal at time $t$ is a Gaussian whose mean shrinks as $\sqrt{\bar\alpha_t}$ and whose variance grows toward one. Push the slider forward and the structured cloud dissolves into noise; a trained reverse network is what walks it back. Slide it and notice that the transformation is a path through distributions, exactly the kind of object the change-of-variables part of the calculus guide studies.

data x encoder q(z|x) → latent z ~ N(μ, σ²) decoder p(x|z) → reconstruction

The forward process adds Gaussian noise to a structured cloud; $t=0$ is data and $t=1$ is nearly pure noise. A trained reverse process learns to run this film backwards.

All three are the same move made three ways. The Gaussian process closes the posterior in a formula; the VAE learns a latent density and a decoder; diffusion learns a chain of conditional denoisers. What they share is that the output of the model is a distribution — and so, for that matter, is the softmax of the next section's classifier. The softmax and the diffusion model differ in the size of the object, not in the kind.

6

Where this shows up

The distribution under every output

AI / Training

The language model's loss

Every language model is a categorical distribution over the vocabulary at each position, trained by the negative log-likelihood of the next token. Perplexity is the exponential of that loss, and the softmax is the layer that defines it.

AI / Evaluation

Calibration and honest confidences

A model whose confidences are trusted downstream must be calibrated, and evaluation measures that with reliability diagrams and ECE. Temperature scaling is the cheap post-hoc fix when the curve misses the diagonal.

AI / Serving

Sampling the next token

The decode loop samples from the temperature-scaled softmax, and speculative decoding compares two such distributions to accept or reject draft tokens. The dial is the same one on this page.

AI / Alignment

Log-probabilities as preferences

Preference training in RLHF scores candidate completions by the log probability the model assigned them — cross-entropy read with a sign flip — and then pushes on the parameters to move that number.

7

Cheat sheet

Every formula in one place

IdeaFormulaReading
Softmax$\sigma(z)_i=e^{z_i}/\sum_j e^{z_j}$Turns logits into a categorical distribution.
Temperature$\sigma(z/T)_i=e^{z_i/T}/\sum_j e^{z_j/T}$Small $T$ sharpens, large $T$ flattens toward uniform.
Entropy$H(p)=-\sum_i p_i\log p_i$How undecided the distribution is, in nats.
Cross-entropy$H(p,q)=-\sum_i p_i\log q_i$Cost of encoding the truth with the model's code.
Negative log-likelihood$-\log q(y)$The cross-entropy of a one-hot label; the training loss.
Decomposition$H(p,q)=H(p)+\mathrm{KL}(p\,\|\,q)$Minimising the loss is minimising the KL divergence.
Perplexity$e^{\mathrm{NLL}}$Effective number of equally likely choices.
Expected calibration error$\sum_b (n_b/N)\,|\mathrm{acc}_b-\mathrm{conf}_b|$Average gap between confidence and accuracy across bins.
Temperature scaling$q=\sigma(z/T)$, $T>1$One parameter that softens an overconfident model.
GP posterior mean$k(x)^\top(K+\sigma_n^2 I)^{-1}y$Kernel-weighted interpolation of the observations.
GP posterior variance$k(x,x)-k(x)^\top(K+\sigma_n^2 I)^{-1}k(x)$Shrinks near data, returns to the prior far away.
RBF kernel$k(x,x')=\sigma_f^2 e^{-(x-x')^2/2\ell^2}$Smoothness set by the lengthscale $\ell$.
Diffusion forward step$q(x_t|x_{t-1})=\mathcal{N}(\sqrt{1-\beta_t}\,x_{t-1},\beta_t I)$Add noise, shrink the signal, keep the variance.
Diffusion marginal$q(x_t|x_0)=\mathcal{N}(\sqrt{\bar\alpha_t}\,x_0,(1-\bar\alpha_t)I)$Closed form for jumping straight to time $t$.
8

Further reading

Where to go deeper

9

Check your understanding

0/6 answered