Likelihood and maximum likelihood
Probability runs forwards: give me a model, and I will tell you how likely a result is. Inference runs the other way. You have already seen the result — a list of numbers, a sequence of coin flips, a corpus of sentences — and you want the model. Which parameter values would have made this particular data unsurprising? That question has a one-line answer, the likelihood, and the answer is not a probability in the usual sense at all. It is a function of the parameter with the data held fixed, and it is exactly the quantity that training a language model, fitting a robot sensor, and estimating an effect size all maximise. This part builds the likelihood, watches it move, and follows it to its peak.
The question
You have the data; which world produced it?
Suppose sixty-four readings come off a sensor. They cluster around some value, with some spread, and the engineering story says the readings are independent draws from a normal distribution whose centre you do not know. The true centre exists — it is a fixed physical quantity — but you never see it. You see only the sample. The forward question, "given the centre, what readings are plausible?", is the probability question you already know how to answer. The inverse question is the one that matters in practice: given these sixty-four numbers, which centres are plausible?
It is tempting to say the centre is just the average of the readings. That is the right answer for this model, but it is a conclusion, not a definition, and it will not survive a change of model. Swap the normal for a Laplace or a Cauchy, or make one of the readings a wild outlier, and the average stops being the natural estimate. What we need is a general principle that takes a model family and a dataset and returns a best-fit parameter, whatever the family is. The principle should reduce to the average where the average is correct, and tell us what to do where it is not.
The principle is to ask, for every candidate parameter value, how well that value would have explained the data we actually got. A model that assigns high probability to the observed sample is a model that had a good chance of producing it. A model that assigns almost no probability to the observed sample is one that calls the data a miracle. Rank the candidate parameters by that score, and take the best. No averaging, no appeal to a prior belief about the parameter, no repeated experiments — just the raw data and the model, compared.
This is a subtle inversion of the direction of explanation. In the forward direction the sample is random and the parameter is fixed. In the inverse direction the sample is fixed — it is the one you have — and we sweep the parameter across its range. The same numbers are involved, but which one varies has changed, and everything downstream depends on keeping that straight. The likelihood is precisely the forward probability, read backwards, with the roles swapped.
The likelihood function
Flip the roles of data and parameter
Write $x = (x_1,\dots,x_n)$ for the observed data and $\theta$ for the parameter, which may be a number or a vector. The model says each observation is a draw from a distribution $p(\cdot;\theta)$, and that the draws are independent and identically distributed. The probability of the whole sample, under that model, is the product of the per-observation probabilities. Now read that product as a function of $\theta$ with x frozen. That is the definition.
The semicolon is doing real work. It separates the argument, $\theta$, from the fixed data, x. If you swap them you get the sampling distribution $p(x;\theta)$, the ordinary probability of the data for a known parameter. The likelihood is that same expression with the emphasis moved. Same formula, different question, and the notation has to keep them apart because the answers are different kinds of thing.
For the Gaussian sensor story, $p(x_i;\mu,\sigma)$ is the bell curve density centred at $\mu$ with width $\sigma$, and the likelihood is the product of sixty-four such densities, one evaluated at each reading. A $\mu$ sitting in the middle of the cloud makes every factor reasonably large. A $\mu$ far off to one side forces many factors towards zero, and their product collapses. The likelihood rewards agreement between the model and the whole sample at once, not any single point.
Independence is what makes the product legitimate, and it is worth noticing how strong a simplification it is. For time series, spatial data, or the tokens of a sentence, observations are not independent, and the product is replaced by a chain of conditional probabilities. The machine-learning story does exactly this: the probability of a document factorises into one conditional distribution per token, given everything before it. The shape of the likelihood changes, but the principle — evaluate the model on the data it actually saw — does not.
Here is the same idea on the simplest possible model, a coin. Flip it twenty times and count k heads. The parameter is the bias p, the probability of heads, and the likelihood is L(p)=p^k(1-p)^{n-k}. Move the slider for the number of heads you observed and watch where the peak of the curve sits. It is always at k/n, the observed frequency: the value of p that makes the data you saw most probable. That peak is the maximum likelihood estimate.
The curve is the likelihood of the coin bias p after seeing k heads in twenty flips. The marker is the MLE, p̂ = k/n, always at the peak.
Two features of that picture recur everywhere. First, the curve is not symmetric in general; at the edges k=0 and k=n it is monotone and its peak sits exactly on the boundary of the parameter space, which is a warning that the "derivative equals zero" recipe has exceptions. Second, the height of the peak is not a probability. It is a density evaluated at one point, and its numerical value depends on how you measured p — as a probability, as an odds, as a log-odds. The peak's location is meaningful; its height is bookkeeping. That distinction is the whole of section four.
The log-likelihood and why we maximise it
Turn a product into a sum, then turn the sum over to calculus
Sixty-four densities multiplied together is a number around 10^{-40}. A million tokens is a number underflowing to zero in double precision long before the product finishes. Multiplying probabilities is numerically hopeless and algebraically awkward: the derivative of a product is a mess of terms. Take the logarithm instead. Logarithm is strictly increasing, so it cannot move the location of a maximum, and it converts the product into a sum. Maximising the product and maximising the sum are the same problem, and the sum is far easier.
The sum is the reason likelihood is practical. Each observation contributes one additive term, $\log p(x_i;\theta)$, so the total is a sum of identical-shaped pieces and its derivatives are sums of derivatives. Under independence, observations combine by addition, which is exactly the structure gradient descent likes. It also ties the likelihood to information: the expected log-likelihood per observation is a cross-entropy, and maximising it is the same as minimising that cross-entropy, which is the loss printed by every training loop.
If the log-likelihood is smooth and the maximum is interior, the derivative at the peak vanishes. That derivative is called the score, and the stationarity condition is the first-order optimality condition you would write for any optimisation problem.
For a concave log-likelihood there is a single peak and the equation $\ell'(\theta)=0$ pins it down. A big family of models has this property: Gaussian means, Bernoulli and multinomial parameters, exponential families in their natural parametrisation. For them the MLE is found by solving the score equation, and a gradient step walks straight uphill. Where the log-likelihood is not concave, the same equation still defines candidate stationary points, but you must check which are maxima; that is why fitting Gaussian mixtures, neural networks, and latent-variable models needs care, and why this volume comes back to optimisation again and again.
The demo below is the heart of this part, so take a moment with it. Three linked views respond to one parameter pair $(\mu,\sigma)$. The top panel superimposes the model density on the histogram of the data, so you can see the fit. The middle panel is the log-likelihood $\ell(\mu)$ with $\sigma$ held at its current value: the marker is where you are, and the dark marker at the peak is the MLE. The bottom panel redraws the identical function as a likelihood on a linear scale, so you can watch the peak sharpen and the tails vanish. Drag the two round handles on the top panel, or the sliders, and every view moves together; press Snap to MLE and the markers collapse onto the peak.
The histogram is the data; the curve is $N(\mu,\sigma)$ at the current parameters. Drag the handles at the curve peak (for $\mu$) and one standard deviation to its right (for $\sigma$).
The log-likelihood $\ell(\mu)=\sum_i\log p(x_i;\mu,\sigma)$, shifted so the peak sits at zero for the current $\sigma$. The marker follows your $\mu$; the dark marker is $\hat\mu=\bar x$.
The same function as a likelihood, scaled to peak at one. Shaded area is $\int L(\mu)\,d\mu$ over the plotted range — the readout insists it is not 1.
For the Gaussian, the score equation solves in closed form. The derivative of $\ell$ with respect to $\mu$ is $\sum_i (x_i-\mu)/\sigma^2$, which is $n(\bar x-\mu)/\sigma^2$ and vanishes exactly at the sample mean. The derivative with respect to $\sigma$ vanishes at the root-mean-square deviation from that mean. So the MLE is the pair that every least-squares fit has been using all along.
Least squares and maximum likelihood are not rivals; for Gaussian noise they are the same calculation, because $\log p(x_i;\mu,\sigma)$ is a constant minus a squared residual over $2\sigma^2$. Maximising the log-likelihood is minimising the sum of squared residuals. Every time a fitting routine reports a residual sum of squares it is reporting a negative log-likelihood up to constants, and every time a language model reports a cross-entropy loss it is reporting a negative log-likelihood per token. The formula is not a textbook curiosity; it is the objective.
Likelihood is not a probability over θ
The most common confusion in statistics, resolved once
Write down $L(\theta;x)$ for a Gaussian with fixed x and sweep $\theta$ over the line. It looks exactly like a bell curve. It is not one. A probability density over $\theta$ must integrate to one; the likelihood has no such obligation, and in general $\int L(\theta)\,d\theta$ is some finite number that depends on the data, the sample size, and the units in which you measured $\theta$. Change the units from metres to centimetres and the curve rescales; the peak stays put, the area does not. A quantity whose total can be anything is not a density.
There is a second, sharper distinction. $L(\theta;x)$ answers "how well does this $\theta$ explain the data?" It does not answer "how likely is this $\theta$?" To get the latter you need something extra: a prior $p(\theta)$ describing what parameter values you considered plausible before seeing the data, and the evidence p(x), which normalises. Bayes' rule then assembles the posterior, and only at that point does a density over $\theta$ exist.
Notice how the likelihood appears unchanged in the numerator. Maximum likelihood is what you get by refusing to supply a prior and reporting the peak alone; Bayesian inference is what you get when you supply one and report the whole reshaped curve. The two are not competing interpretations of the same curve so much as two different uses of it, and the next part of this volume is entirely about the second use.
The demo below isolates the point about the integral. It shows a one-parameter likelihood built from n observations, scaled so its peak is one, with the area under it shaded. As you raise n, the curve narrows like $1/\sqrt n$ and the area shrinks with it. Nothing is wrong: a function that peaks at one and integrates to a tenth is a perfectly good likelihood. It simply cannot be a probability density over the parameter, because it does not carry the total probability that a density must.
A likelihood over $\mu$ from n observations, scaled to peak at one. The shaded region is the area the readout reports; watch it fall as n grows.
One more consequence is worth naming because it confuses people constantly: the likelihood is a function of the data. Two experiments that produce different samples give different likelihood curves, and you cannot compare the heights of curves from different datasets as if they were probabilities on a common scale. A likelihood ratio, however — the same curve divided by its own maximum, or two nested models compared on identical data — is invariant to constants and units, and that ratio is what tests, confidence intervals, and the hard part of Bayesian model comparison are built from.
Properties and asymptotics
Why the peak is trustworthy when there is enough data
The maximum likelihood estimate has a short list of guarantees, and they explain why it is the default. Under mild regularity conditions — a smooth, identifiable model whose true parameter sits in the interior of the parameter space — the estimate is consistent: as the sample grows, it converges to the true parameter in probability. The reason is visible in the curve. Each observation contributes one term to $\ell$, and by the law of large numbers the average term converges to the expected log-likelihood, which is maximised at the truth. Adding data sharpens the peak around that point, so the estimate is dragged to the right answer.
It is also asymptotically normal and efficient. The error $\hat\theta-\theta$ is approximately Gaussian with variance set by the curvature of the log-likelihood at the peak, and no unbiased estimator can do better in the large-sample limit. The curvature has a name: the Fisher information, the negative expected second derivative of the log-likelihood. A sharp peak means large information, a precise estimate, and small standard error; a flat peak means the data barely distinguish nearby parameters. This is the bridge from likelihood to the intervals, error bars, and information bounds of the next parts.
Two practical caveats sit inside that clean story. First, the estimator can be biased in finite samples even when it is consistent, and the Gaussian spread is the standard example: dividing by n systematically underestimates $\sigma^2$, and the familiar n-1 correction is the unbiased repair. Second, the guarantees are asymptotic, and the approximation degrades when the sample is small, the parameter sits on a boundary, or the model is unidentifiable because two different parameters produce the same distribution. A mixture of two Gaussians is the canonical trap: swapping the two components gives an identical likelihood, so the peak is a ridge, not a point.
There is one more property that makes likelihood convenient in modelling rather than just estimation. The MLE is invariant under reparameterisation: if you find the peak in terms of a standard deviation, then the peak in terms of a variance is just the square of it. You never have to redo the optimisation because you changed how you write the parameter. This is not true of unbiasedness — a median-unbiased estimate of one quantity can be biased for a function of it — and it is one reason likelihood became the organising principle of modern statistics rather than just one method among many.
Finally, the likelihood is the natural language for comparing models of different complexity. Add parameters and the peak can only rise, because a bigger family contains the smaller one; that is why a raw likelihood cannot rank a simple model against a complicated one. Penalise the extra freedom and you get the information criteria, and hold part of the data aside and you get cross-validation. Both are attempts to ask not "which model fits?" but "which model would predict new data?" — a question this volume returns to whenever it fits something.
Where this shows up
One objective, many names
Pretraining is maximum likelihood
A language model assigns a probability to the next token given the context, and pretraining maximises the log-likelihood of the corpus, which the training loop reports as cross-entropy loss per token. Everything in this part applies with the model's weights playing the role of $\theta$; the gradient of the loss is the negative score.
Scaling laws are loss curves
Empirical scaling laws fit smooth functions to how the negative log-likelihood falls with compute and data. They are likelihood curves read across model sizes, and the irreducible floor they extrapolate to is the entropy of the data itself.
Optimisers climb the log-likelihood
Gradient ascent on $\ell$ is gradient descent on the negative log-likelihood, and optimizers like momentum and Adam are the machinery that reaches the peak when the score equation has no closed form. The demo's Snap to MLE is what those methods do, slowly and in high dimensions.
Densities and change of variables
The likelihood is a density, and pushing a random variable through a map rescales it by a Jacobian. That is the calculus of probability covered in Calculus of probability, and it is why a maximum likelihood problem in transformed coordinates is a different optimisation problem even though the answer transforms cleanly.
Cheat sheet
Every formula in one place
| Idea | Formula | Intuition |
|---|---|---|
| Likelihood | $L(\theta;x)=\prod_i p(x_i;\theta)$ | Product of per-observation probabilities, read as a function of $\theta$. |
| Log-likelihood | $\ell(\theta)=\sum_i\log p(x_i;\theta)$ | Same peak, additive terms, no underflow; the numerical workhorse. |
| MLE | $\hat\theta=\arg\max_\theta \ell(\theta;x)$ | The parameter that makes the observed data least surprising. |
| Score | $\ell'(\theta)=\sum_i \partial_\theta\log p(x_i;\theta)=0$ | Stationarity condition; the peak of a concave log-likelihood. |
| Gaussian MLE | $\hat\mu=\bar x,\quad \hat\sigma^2=\frac1n\sum_i (x_i-\bar x)^2$ | Sample mean and mean squared deviation; least squares in disguise. |
| Relative likelihood | $L(\theta;x)/L(\hat\theta;x)\in[0,1]$ | Invariant to constants and units; the basis of likelihood-ratio tests. |
| Not a density | $\int L(\theta;x)\,d\theta\ne1$ | Only prior times likelihood over evidence makes a posterior density. |
| Fisher information | $I(\theta)=-\mathbb{E}[\ell''(\theta)]$ | Curvature at the peak; inverse variance of the asymptotic estimate. |
| Posterior | $p(\theta\mid x)\propto L(\theta;x)\,p(\theta)$ | The likelihood seen through the lens of a prior. |
Further reading
Where to go deeper
- R. A. Fisher, On the mathematical foundations of theoretical statistics, Philosophical Transactions of the Royal Society A, 1922 — the paper that introduced likelihood and maximum likelihood estimation.
- A. W. F. Edwards, Likelihood, expanded edition, 1992 — the case for treating likelihood as the central concept of inference, with the geometry of the surface.
- David MacKay, Information Theory, Inference, and Learning Algorithms, chapters 2 and 22 — likelihood, the Gaussian MLE, and the jump to Bayesian inference, written for this audience.
- Lucien Le Cam, Asymptotic Methods in Statistical Decision Theory, 1986 — the rigorous version of the consistency and asymptotic-normality story in section five.
- Bradley Efron and Trevor Hastie, Computer Age Statistical Inference, chapters 4–5 — maximum likelihood alongside the resampling ideas of the next parts of this volume.