Probability in machine learning
Strip the marketing away and a modern model is a probability distribution with a large parameter count. A classifier outputs a categorical distribution over labels; a language model outputs one over the next token; a diffusion model learns to undo a noising process that is itself a chain of conditional Gaussians. Training is maximum likelihood in disguise, and every loss you have seen — cross-entropy, log loss, denoising score matching — is a negative log-likelihood of something. The subject of this volume was built for exactly these objects, and this last part walks the loop: a distribution is produced by a network, scored against reality by a likelihood, and then asked how much it should be trusted. By the end the softmax, the loss curve, the reliability diagram and the uncertainty band are not separate tricks but four views of one idea.
The question
Where does the probability actually live?
Ask a working engineer where probability sits inside a neural network and the answer usually points at the loss. That is half right. The loss is where probability is used, but the probability is produced earlier, by an output layer that turns a vector of real numbers into something that sums to one. For a classification head that layer is the softmax; for a language model it is the same softmax applied to a vocabulary-sized vector of logits; for a Gaussian head it is a mean and a variance. Every one of those is a distribution parameterised by the network, and the network's job is to move that distribution onto the data.
Once you see the output as a distribution, the training objective stops being arbitrary. The likelihood of the observed label under the predicted distribution is a number between zero and one; training maximises it. Maximising a product is the same as minimising the sum of negative logs, which is why the loss is written with a minus sign and a logarithm. Cross-entropy is not a separate invention bolted onto the softmax — it is the negative log-likelihood of the categorical model the softmax defines. Flip the question around and the same identity explains why the loss is large when the model is confidently wrong: the predicted probability of the true label was tiny, and the negative log of a tiny number is large.
Two further questions follow from that picture, and they are where the probabilistic reading earns its keep. First, is the number the model reports true? A network that says 0.9 should be right about ninety percent of the time on that kind of input, and if it is not, the distribution is miscalibrated even when the accuracy is fine. Calibration is a statement about the distribution, not the decision, and it has its own diagnostic — the reliability diagram — and its own cheap fix, temperature scaling. Second, how much should a single prediction be believed at all? A point estimate throws away the shape of the distribution. Gaussian processes, variational autoencoders and diffusion models each preserve it in a different way: a closed-form posterior, a learned latent code, or a learned reverse process.
This part is the volume's cashing-in chapter. The softmax is a distribution built in the information theory part; the likelihood is Part 1; the latent-variable ELBO is Part 12; the Gaussian conditioning behind the GP is the multivariate Gaussian part of Volume I. Nothing new is introduced except the assembly. The plan is to take the output distribution apart, then the loss, then the calibration of the loss, then uncertainty, and finally the generative models that make the whole distribution the object being learned.
Softmax as a distribution
A dial that controls how committed the model is
A network's final linear layer hands you a vector $z$ of real numbers called logits. They have no units and can be any size or sign; the only thing that matters is their differences, because adding a constant to every logit changes nothing about the ranking. To turn them into probabilities you exponentiate and normalise. The exponential makes every entry positive and makes differences multiplicative, and the sum in the denominator forces the result to add to one. That is the whole definition:
The second expression introduces a temperature $T$. Dividing every logit by $T$ before the softmax reshapes the distribution without changing which class wins. When $T$ is small the logits are stretched apart, the largest entry dominates, and the distribution becomes nearly one-hot — the model is confident. When $T$ is large the logits are squashed together and the distribution flattens toward the uniform $1/K$, which has the maximum entropy any $K$-class distribution can have. This one dial is the difference between a greedy sampler and a creative one.
You have already met the dial in the serving guide. The decode loop samples each token from the softmax over the next-token logits, and the temperature is the knob that trades diversity for reliability. Set it to zero and decoding becomes greedy argmax; raise it and the tail of the distribution gets sampled. Speculative decoding takes the same distribution and uses a small draft model to propose tokens that a larger model then accepts or rejects, and the accept-reject test is written in terms of these probabilities — it is a couple of distributions being compared, nothing more.
Drag the four logits and the temperature. Two things are worth watching. The ranking of the bars is fixed by the logits; the temperature only changes how spread out they are. And as $T$ falls the entropy drops toward zero, while as $T$ rises it climbs toward the uniform ceiling $\log K$. The distribution is doing one smooth thing and the readout names it.
Bars are the class probabilities $\operatorname{softmax}(z/T)$. The dashed line is the uniform distribution $1/K$, the flattest this softmax can get.
The distribution the softmax defines is called categorical, and its entropy is the number of nats of information the model is undecided by. That is not a metaphor: a distribution with entropy $H$ costs at least $H$ nats to encode, so an uncertain model is literally an expensive one. The next section makes that accounting exact by scoring the distribution against what actually happened.
Cross-entropy is negative log-likelihood
The loss is just the log of the probability you assigned to the truth
Suppose the true class is $y$ and the model predicted the distribution $q$. The likelihood of the observation under the model is $q(y)$, a single number. Training wants that number to be as large as possible, and because logs turn products into sums, the equivalent objective is to minimise its negative logarithm. One example, one term:
The middle expression is the cross-entropy between a true distribution $p$ and the model's $q$. When the label is a one-hot vector — all its mass on the true class — the sum collapses to the single term $-\log q(y)$, which is the negative log-likelihood. The identity on the right is the reason cross-entropy is the standard loss: it equals the entropy of the truth, which the model cannot change, plus the KL divergence from truth to model. Minimising cross-entropy therefore minimises the KL divergence, the same quantity the information theory part read as an excess codelength. And by Gibbs' inequality the KL is never negative, so the loss bottoms out at the true entropy and no lower.
The shape of the function carries the practical lesson. Because the logarithm is steep near zero, a confident mistake is punished hard: predicting $q(y)=0.01$ costs about $4.6$ nats, while $q(y)=0.5$ costs only $0.69$. A model that puts almost no mass on the truth is driven by an enormous gradient, which is exactly the pressure that makes a classifier stop being confidently wrong. It also explains why numerical implementations clamp the log or use a fused log-softmax: the raw logarithm of a probability that underflowed to zero is not a number you want to compute.
In the demo the model's predicted distribution is fixed to one shape — a large logit on class 0, zeros elsewhere — and you choose which class is the truth. When the truth is class 0 the loss is small and falls as the margin grows. When the truth is any other class the model is confidently wrong, $q(y)$ is minuscule, and the loss explodes. Move the margin slider with the truth set away from the peak and watch the meter fill.
Bars are the predicted distribution $q$. The dashed line marks the true class $y$; the loss is $-\log q(y)$, read from the height of the bar the line sits on.
This is the same object the language-model guide calls the training loss and the evaluation guide calls perplexity. Perplexity is just $e^{\mathrm{NLL}}$, the effective number of equally likely choices the model was entertaining; a perplexity of $K$ means the model was as confused as a uniform distribution over $K$ tokens. And preference training in the RLHF part starts from the log probabilities the same softmax produced, so the whole pipeline from pretraining to alignment is arithmetic on a categorical distribution.
Calibration and reliability
A probability is a promise about frequency
A classifier that outputs $0.9$ is making a falsifiable claim: on all the inputs it labels with confidence near $0.9$, about ninety percent should be correct. A model that satisfies this is calibrated. Calibration is separate from accuracy. A model can be accurate and wildly overconfident (right on most inputs, but claiming $0.99$ when it is right only $0.8$ of the time), or accurate and underconfident. Modern deep networks are usually the first: they win on accuracy and then report confidences no one should trust.
The diagnostic is a reliability diagram. Bin the predictions by their stated confidence, then for each bin compute the empirical accuracy — the fraction of those predictions that were actually correct. Plot accuracy against confidence. A perfectly calibrated model lies on the diagonal $y=x$, because its claimed confidence equals the observed frequency. Points below the diagonal are overconfident; points above are underconfident. The gap between the curve and the diagonal, weighted by how many predictions fall in each bin, is the expected calibration error:
Overconfidence has a remarkably cheap repair called temperature scaling. Leave the trained network alone and divide its logits by a single number $T>1$ before the softmax. This softens every distribution, pulling the confidence down toward the accuracy without changing a single predicted class, because dividing by a positive number does not reorder the logits. You fit the one parameter $T$ on a held-out set by minimising the same negative log-likelihood, and the reliability curve walks back to the diagonal. One scalar repairs a network with millions of weights.
The demo simulates a classifier whose predictions are warped by an overconfidence factor. The true probability of the positive class is drawn first, the label is sampled from it, and the model's confidence is a sharpened version of the truth, so its errors are exactly the overconfidence you would see in practice. The pink points are the raw model; the blue points are the same model after temperature scaling. Slide the temperature and watch the blue curve snap onto the diagonal while the class decisions never move.
Reliability diagram: predicted confidence on the $x$-axis, empirical accuracy on the $y$-axis. Pink is the raw model, blue is temperature-scaled; the pink verticals are the calibration gaps that add up to the ECE.
Calibration is not the same as a good decision rule. If all you need is the argmax, a miscalibrated model may serve perfectly well; if the probability is going to be multiplied by a cost, fed into a Bayesian filter, or used to decide whether to hand a task to a human, then the number itself carries weight and ECE is the thing to watch. This is also why evaluation sets double as calibration sets in the evaluation part: the same held-out data that measures accuracy measures the honesty of the confidences.
Uncertainty: GPs, VAEs, diffusion
Keeping the whole distribution instead of one point
A point prediction throws the distribution away at the last moment. Three families of model refuse to. The first is the Gaussian process, which puts a prior over functions: any finite set of function values is jointly Gaussian, and conditioning on observations gives a posterior that is itself Gaussian everywhere. With an RBF kernel and a handful of noisy observations, the posterior mean is a smooth interpolation and the posterior variance has a shape you can read: it collapses toward the noise floor at each observed input and swells back toward the prior far away from all of them. The formulas are the multivariate Gaussian conditioning of Volume I, written for an infinite index set:
Drag the lengthscale and the noise and watch the band. A long lengthscale makes the function smooth and the uncertainty broad; a short one lets the posterior wiggle and the band pinch tightly at every observation. More noise raises the floor and keeps the band from ever reaching zero, because the observations themselves are not trusted. The important thing is not the curve — it is the second line of the plot, the variance, which no point-estimate model can draw at all.
The line is the posterior mean; the shaded band is $\pm 2\sigma_*$. Markers are the noisy observations. The band pinches at each marker and widens back toward the prior between them.
The second family is the variational autoencoder, which buys a distribution over latent codes by learning the transformation itself. An encoder maps a data point $x$ to a Gaussian $q(z\mid x)$ in a low-dimensional latent space, a decoder maps a code $z$ back to a distribution over data, and training maximises the evidence lower bound from Part 12: a reconstruction term plus a KL term that keeps the codes near a standard normal. Sampling from that latent Gaussian and decoding gives a whole cloud of plausible outputs, and the KL term is what makes the latent space smooth enough for interpolation to mean something.
The third family is diffusion, which learns a distribution by learning to destroy and then rebuild structure. The forward process is a fixed chain of Gaussians that adds a little noise at each step, and the reverse process is a network trained to undo one step at a time. The forward step is the simple object, and once you see it the rest of the pipeline is a matter of parameterisation:
Each step keeps the signal coefficient $\sqrt{1-\beta_t}$ and adds variance $\beta_t$, so the chain is variance-preserving and the marginal at time $t$ is a Gaussian whose mean shrinks as $\sqrt{\bar\alpha_t}$ and whose variance grows toward one. Push the slider forward and the structured cloud dissolves into noise; a trained reverse network is what walks it back. Slide it and notice that the transformation is a path through distributions, exactly the kind of object the change-of-variables part of the calculus guide studies.
The forward process adds Gaussian noise to a structured cloud; $t=0$ is data and $t=1$ is nearly pure noise. A trained reverse process learns to run this film backwards.
All three are the same move made three ways. The Gaussian process closes the posterior in a formula; the VAE learns a latent density and a decoder; diffusion learns a chain of conditional denoisers. What they share is that the output of the model is a distribution — and so, for that matter, is the softmax of the next section's classifier. The softmax and the diffusion model differ in the size of the object, not in the kind.
Where this shows up
The distribution under every output
The language model's loss
Every language model is a categorical distribution over the vocabulary at each position, trained by the negative log-likelihood of the next token. Perplexity is the exponential of that loss, and the softmax is the layer that defines it.
Calibration and honest confidences
A model whose confidences are trusted downstream must be calibrated, and evaluation measures that with reliability diagrams and ECE. Temperature scaling is the cheap post-hoc fix when the curve misses the diagonal.
Sampling the next token
The decode loop samples from the temperature-scaled softmax, and speculative decoding compares two such distributions to accept or reject draft tokens. The dial is the same one on this page.
Log-probabilities as preferences
Preference training in RLHF scores candidate completions by the log probability the model assigned them — cross-entropy read with a sign flip — and then pushes on the parameters to move that number.
Cheat sheet
Every formula in one place
| Idea | Formula | Reading |
|---|---|---|
| Softmax | $\sigma(z)_i=e^{z_i}/\sum_j e^{z_j}$ | Turns logits into a categorical distribution. |
| Temperature | $\sigma(z/T)_i=e^{z_i/T}/\sum_j e^{z_j/T}$ | Small $T$ sharpens, large $T$ flattens toward uniform. |
| Entropy | $H(p)=-\sum_i p_i\log p_i$ | How undecided the distribution is, in nats. |
| Cross-entropy | $H(p,q)=-\sum_i p_i\log q_i$ | Cost of encoding the truth with the model's code. |
| Negative log-likelihood | $-\log q(y)$ | The cross-entropy of a one-hot label; the training loss. |
| Decomposition | $H(p,q)=H(p)+\mathrm{KL}(p\,\|\,q)$ | Minimising the loss is minimising the KL divergence. |
| Perplexity | $e^{\mathrm{NLL}}$ | Effective number of equally likely choices. |
| Expected calibration error | $\sum_b (n_b/N)\,|\mathrm{acc}_b-\mathrm{conf}_b|$ | Average gap between confidence and accuracy across bins. |
| Temperature scaling | $q=\sigma(z/T)$, $T>1$ | One parameter that softens an overconfident model. |
| GP posterior mean | $k(x)^\top(K+\sigma_n^2 I)^{-1}y$ | Kernel-weighted interpolation of the observations. |
| GP posterior variance | $k(x,x)-k(x)^\top(K+\sigma_n^2 I)^{-1}k(x)$ | Shrinks near data, returns to the prior far away. |
| RBF kernel | $k(x,x')=\sigma_f^2 e^{-(x-x')^2/2\ell^2}$ | Smoothness set by the lengthscale $\ell$. |
| Diffusion forward step | $q(x_t|x_{t-1})=\mathcal{N}(\sqrt{1-\beta_t}\,x_{t-1},\beta_t I)$ | Add noise, shrink the signal, keep the variance. |
| Diffusion marginal | $q(x_t|x_0)=\mathcal{N}(\sqrt{\bar\alpha_t}\,x_0,(1-\bar\alpha_t)I)$ | Closed form for jumping straight to time $t$. |
Further reading
Where to go deeper
- Christopher Bishop, Pattern Recognition and Machine Learning, chapters 4 and 6 — the logistic softmax and the Gaussian process regression chapter this part compresses.
- Carl Rasmussen and Christopher Williams, Gaussian Processes for Machine Learning, 2006 — the definitive treatment, free online, with the kernel and conditioning formulas used here.
- Guo, Pleiss, Sun and Weinberger, "On Calibration of Modern Neural Networks", 2017 — the reliability diagrams and temperature scaling result that made calibration a standard topic.
- Kingma and Welling, "Auto-Encoding Variational Bayes", 2013 — the VAE and the ELBO, the latent-variable construction of Part 12 applied at scale.
- Ho, Jain and Abbeel, "Denoising Diffusion Probabilistic Models", 2020 — the forward process written down and the reverse process trained to undo it.
- Ian Goodfellow, Yoshua Bengio and Aaron Courville, Deep Learning, chapters 5 and 6 — maximum likelihood, cross-entropy and the probabilistic view of network outputs.