Maximum likelihood
A sample fixes the numbers and leaves the distribution unknown. Maximum likelihood turns that asymmetry around: hold the data still, treat the parameters as the variable, and ask which settings make the observed numbers look least surprising. This part makes the question physical. The data sit as ticks on an axis and a normal density hovers over them; drag its mean and its spread and a live log-likelihood surface shows how plausible your parameters are. The peak of that surface is the maximum-likelihood estimate — the fit that puts the density squarely onto the data.
The question
Which parameters make the data look likely?
Probability and statistics point in opposite directions. In probability the distribution is known and the data are random: you are told a coin is fair and asked what runs it can produce. In statistics the data are known and the distribution is not: you are handed one run of flips and asked what the coin's bias might be. Parts 1–6 estimated moments without committing to any particular shape for the population. Maximum likelihood goes further and asks you to name a family of distributions, a formula $p(x;\theta)$ with an unknown parameter $\theta$, and then chooses the member of that family under which the observed data would have been most probable.
That one sentence hides the whole method, so the demo in this part slows it down until the idea is visible. A fixed sample of numbers is drawn once and then frozen. The parameter is the only thing that moves. As you drag it, the density that the parameter describes slides across the data, and a surface records how good the fit is at every setting. The peak of the surface is the answer.
The likelihood: read the density backwards
Same formula, new argument
Suppose the sample is $x_1, \dots, x_n$ and each observation is drawn independently from the same density $p(x;\theta)$. Independence lets the joint density of the whole sample factorise, and identical distribution lets one symbol $\theta$ serve them all:
Nothing has happened yet; this is just the probability of seeing this sample when the parameter is $\theta$. The move that creates statistics is to forget that the $x_i$ were random and to read the very same expression as a function of $\theta$, with the observed values plugged in and held constant. That function is the likelihood:
It answers a subtly different question. The density told you how probable each possible data value was for a given $\theta$; the likelihood tells you how plausible each possible $\theta$ is for the one data value you actually observed. The words are similar and the picture is completely different — the same curve, evaluated along a different axis. This is the same flip that the probability chapter makes when it moves from a density over outcomes to a weight over hypotheses.
Multiplying many small numbers is numerically fragile and analytically unpleasant, and the logarithm fixes both problems without changing the answer. Because $\log$ is strictly increasing, the $\theta$ that maximises $L$ also maximises
a sum of log densities instead of a product. The log-likelihood is the object every optimizer and error bar in this series is built from: whenever a likelihood is multiplied, its logarithm is added.
Drag the parameters
A density over the data, a surface over the parameters
The demo freezes one sample of forty values, drawn from a normal distribution. The ticks along the baseline of the lower canvas are those values. The solid curve is a normal density at the current parameters $(\mu,\sigma)$; the two round handles on the curve control it. Drag the $\mu$ handle left and right and the whole density slides along the axis, following the handle. Drag the $\sigma$ handle horizontally and the density widens or narrows, keeping its total area at one. The dashed curve is the best fit, found by the formulas of Step 5, and it stays put as a reference.
Above the canvas is the signature picture of this part: the log-likelihood surface $\ell(\mu,\sigma)$ drawn as a filled contour plot over a grid of candidate parameters. Brighter means a higher log-likelihood, so brighter means a better explanation of the frozen data. The code recomputes $\ell$ from scratch at every grid point using the definition $\ell(\mu,\sigma) = \sum_i \log \mathcal{N}(x_i;\mu,\sigma)$. The marker sitting on the surface is your current $(\mu,\sigma)$. Slide the density's handles and the marker moves with them; the readout reports the gap between your log-likelihood and the best one.
Ticks are the data; the solid curve is the density at your $(\mu,\sigma)$; the dashed curve is the maximum-likelihood fit. Drag the round handles.
Log-likelihood surface $\ell(\mu,\sigma)$ over a grid: brighter is higher. The marker is your parameters, and the bright peak is the MLE.
Why the maximum
The score, and the equation it solves
The maximum-likelihood estimate (MLE) is the parameter that maximises the likelihood, $\hat{\theta} = \arg\max_\theta \ell(\theta)$. Maximising a smooth function is a calculus problem, and the answer is usually found by differentiating. The derivative of the log-likelihood has a name of its own, the score:
For a one-dimensional parameter we look for the point where the slope of the log-likelihood is zero. In several parameters $\theta = (\theta_1,\dots,\theta_k)$ the score is a vector of partial derivatives and the condition is that every component vanish. This is the score equation, and for the normal model it can be solved in closed form, which is why the answer is a formula rather than a numerical search.
Two properties of the log-likelihood are what make the search tractable. First, it is additive across observations, so each data point contributes its own term $\log p(x_i;\theta)$ and the score is a sum of independent contributions. Second, the curve is concave for many standard models, the normal among them: it rises to a single peak and falls away in every direction, so a zero of the score is the global maximum and not merely a local one. The contour plot in the demo shows exactly this shape — one bright basin, no other peaks to fall into.
When there is no closed form, the same score equation is solved numerically by walking uphill, which is the subject of the optimizers chapter and of gradient descent. The curvature of the peak is not an afterthought either: the second derivative of $\ell$, the Hessian, measures how sharply the likelihood falls away from its maximum, and it will become the standard error and then the Fisher information of the next part. The Jacobian and Hessian chapter is the calculus underneath that. The readout in the demo already reports a preview: observed-information standard errors at your current $(\mu,\sigma)$.
It is worth saying what the likelihood is not doing. It does not claim the MLE is the true parameter, only that it is the member of the chosen family under which the observed data are least surprising. If the family is wrong — if you model skewed data with a normal — the maximum is still well defined and still wrong in a systematic way.
The normal MLE
Two formulas you can derive in a page
The normal family is the one used everywhere else in this series, and its MLE is the cleanest possible illustration. Write the density and collect the sample into a single expression. For $x_i \sim \mathcal{N}(\mu,\sigma^2)$ independently, the log-likelihood is
Differentiate with respect to $\mu$, holding $\sigma^2$ fixed. Only the last term depends on $\mu$, and setting its derivative to zero gives $\sum_i (x_i-\mu) = 0$, whose solution is the sample mean:
So the most familiar summary statistic in the world is also the maximum-likelihood estimate of the normal mean. That is not a coincidence to be memorised; it falls out of the score equation in two lines. Now differentiate with respect to the variance and set the result to zero. The derivative is $-n/(2\sigma^2) + (1/2\sigma^4)\sum_i(x_i-\hat{\mu})^2$, and solving it gives the variance estimate
Note the divisor: it is $n$, not $n-1$. The MLE divides the sum of squared deviations by the number of observations. The difference is small when $n$ is large and matters when it is not: the MLE of the variance is biased downward, because it measures spread about the estimated mean $\bar{x}$ rather than about the unknown true mean, and the fitted mean is always a little closer to the data than the truth is. The unbiased estimator divides by $n-1$ instead. Neither is simply "right". The unbiased rule is built to have the correct average over repeated samples; the MLE is built to be the most likely explanation of the sample in front of you. This series reaches for each when its purpose calls for it, and Part 6 is where the trade-off between them is laid out.
The slope of the log-likelihood at the peak also gives an error bar almost for free. For the normal mean the observed information is $\hat{I}(\mu) = n/\hat{\sigma}^2$, and inverting it gives an approximate standard error $\hat{\sigma}/\sqrt{n}$ — the same $\mathrm{SE}$ the earlier parts derived from the sampling distribution, now recovered from the curvature of a single likelihood. The demo prints this quantity as you drag, and the next part builds the whole theory of bounds on top of it.
Not a probability, and it travels well
Two misreadings to avoid, one property to keep
The first misreading is to treat $L(\theta)$ as a distribution over the parameters. It is not. A probability density over $\theta$ would have to integrate to one over the parameter space, and the likelihood generally does not; its integral over $\theta$ has no meaning, and its value at a point is not "the probability that $\theta$ is correct". What is true is only that relative heights carry information: a parameter with twice the likelihood is twice as plausible under this model, and a parameter at half the likelihood is half as plausible. Multiplying the entire likelihood by any constant, including a constant that depends on the data, leaves the peak exactly where it was. That is why dropping the $-\tfrac{n}{2}\log(2\pi)$ term changed nothing. The Bayesian part of this series, later on, does turn the likelihood into a distribution over $\theta$ — but only by multiplying it by a prior and dividing by the total, which is the extra step that normalises it into a posterior. Until then, resist the temptation to read areas under the likelihood curve.
The second good property is invariance. If $\hat{\theta}$ maximises the likelihood and $\phi = g(\theta)$ is any transformation of the parameter, then the maximum-likelihood estimate of $\phi$ is simply $g(\hat{\theta})$. Maximise first and transform after, or reparameterise first and maximise the new likelihood, and you land in the same place. The demo uses this without ceremony: the variance MLE is $\hat{\sigma}^2$, so the standard-deviation MLE is just its square root, $\hat{\sigma} = \sqrt{\hat{\sigma}^2}$. The surface $\ell$ drawn against $\sigma$ rather than $\sigma^2$ looks different — squashed and reshaped — but its ridge sits over the same data summary. This is a property few estimators enjoy, and it is a large part of why the likelihood is the default starting point for almost every modern model, from regression to the deep networks of the ML parts of this site.
Put together, the likelihood is a ranking device rather than a probability: it says which parameter is better and by what ratio, and that ranking survives reparameterisation, so you can report the estimate on whichever scale is most useful.
Where this shows up
One principle, two very different machines
Least squares is Gaussian likelihood
Fitting a line by minimising the sum of squared residuals is not a separate idea from maximum likelihood; it is maximum likelihood for a normal error model. If each observation is $y_i \sim \mathcal{N}(a + bx_i, \sigma^2)$ with the same $\sigma$, the log-likelihood is largest exactly where the sum of squared residuals is smallest, and the score equations are the normal equations. The least-squares part solves the same problem by projection, so its algebra and this part's probability are two views of one fit.
MLE in SLAM
Pose-graph optimisation, the engine of modern SLAM, is maximum likelihood written large: choose the robot poses that best explain a set of relative measurements, each with its own Gaussian noise. The result is a weighted least-squares problem whose Hessian is the information matrix, and the optimiser is a direct descendant of walking up the log-likelihood surface shown here. The pose-graph part makes that connection explicit; the same principle drives the calibration and triangulation problems in the multi-view geometry guide.
The principle repeats wherever a model has parameters and data. Pre-training a language model maximises a (log-)likelihood of tokens under the model, which is why the training loss is reported as a negative log-likelihood per token; the pre-training chapter and the evaluation chapter are reading the same quantity the score equation defines, on a scale measured in billions of terms. Logistic regression, mixture models, and the noise models behind Kalman filters are all likelihoods differentiated and maximised. Once you can see the log-likelihood as a surface and the estimator as its summit, the formulas that follow read as geometry you can point at.
Further reading
The references below develop maximum likelihood as the organising principle of estimation rather than a formula to apply. The one idea to keep is the picture of a log-likelihood surface and the parameter that sits at its peak.
Wasserman treats likelihood inference in a single compact chapter and is the treatment this series follows in spirit; Casella and Berger derive the score equation, the invariance property and the asymptotic normality of the MLE at the level a first graduate course expects. The 1922 Fisher paper is the origin of the method and is surprisingly readable. For the intuition, the visual walkthroughs are the fastest way to make the surface feel concrete.
- Larry Wasserman, All of Statistics: A Concise Course in Statistical Inference, chapter 9 — parametric inference and maximum likelihood.
- George Casella and Roger Berger, Statistical Inference, chapter 7 — the score function, the MLE and its invariance.
- R. A. Fisher, "On the mathematical foundations of theoretical statistics", 1922 — where likelihood was introduced as a principled criterion.
- StatQuest, "Maximum Likelihood, clearly explained" — the same slide-the-density picture, worked through on screen.
Cheat sheet
| Term | Meaning here |
|---|---|
| Likelihood $L(\theta)$ | The joint density of the observed data read as a function of $\theta$ |
| Log-likelihood $\ell(\theta)$ | $\sum_i \log p(x_i;\theta)$; same maximiser as $L$, far easier to handle |
| MLE $\hat{\theta}$ | The parameter that maximises $L$ or, equivalently, $\ell$ |
| Score $U(\theta)$ | The derivative $d\ell/d\theta$; the MLE sets the score to zero |
| Normal $\hat{\mu}$ | The sample mean $\bar{x}$ |
| Normal $\hat{\sigma}^2$ | The divide-by-$n$ variance; biased low, unlike the divide-by-$(n-1)$ rule |
| Invariance | The MLE of $g(\theta)$ is $g(\hat{\theta})$ for any transformation $g$ |
| Not a probability over $\theta$ | Only relative heights matter; normalising it into a posterior needs a prior, in the Bayesian part of the series |