Fisher information and the Cramér–Rao bound
The curvature of the log-likelihood is the error bar. A sharply peaked log-likelihood means the data pin the parameter down tightly; a flat one means they barely constrain it at all. Fisher information is the number that makes that sentence precise, the Cramér–Rao bound turns it into a floor under the variance of every unbiased estimator, and the maximum-likelihood estimate is the one rule that reaches that floor. This part lets you zoom the peak, read off its width, and watch the sampling distribution of the estimate narrow onto the bound as the sample grows.
The question
How much does a sample know about a parameter?
Part 7 fitted a parameter by finding the top of a log-likelihood curve. The top tells you where the best guess is; it says nothing about how sharp that best guess is. A log-likelihood can be a narrow spike, in which case the data nearly determine the parameter, or a broad dome, in which case a wide range of values fits almost as well. The information in the sample is exactly that sharpness, and the whole of this part is one idea: the curvature of the log-likelihood at its peak measures how much the data know.
This is not a metaphor that happens to be useful. It is a theorem with a physical reading. Differentiate the log-likelihood once and you get the score, the direction and steepness of the slope. Differentiate again and you get the curvature, which tells you how quickly the slope falls away from the peak. The average squared score, and the average curvature, are the same number — the Fisher information — and that number sets a hard limit on how precisely any honest estimator can report the parameter.
Two questions organise the rest. First, what is the information, concretely, for a model simple enough to compute it by hand? Second, what does it forbid? The answer to the second is the Cramér–Rao lower bound: no unbiased estimator can beat the reciprocal of the information, and maximum likelihood, in the limit of large samples, exactly attains it. From there, experimental design becomes an information-accounting problem: to shrink an error bar you must add information, and adding data is only one of the ways to do it.
The score function
The slope of the log-likelihood
Fix a model $p(x \mid \theta)$ with one unknown parameter $\theta$, and draw one observation $x$. The log-likelihood of that single observation is $\ell(\theta) = \log p(x \mid \theta)$, and its derivative
is the score. It measures how sensitively the log-probability of what you saw responds to a nudge of the parameter. At the maximum-likelihood estimate the score of the whole sample is zero — the peak is level — and the score of a single observation tells you which way that observation is pulling.
The score has a property that looks accidental and is not. Its expectation under the model is zero:
The score is a random variable with mean zero, so all of its interesting content is in its variance. That variance is the Fisher information of a single observation,
and for $n$ independent observations the scores add, so their variances add too and the information is $n\,I(\theta)$. This is the first place the famous square-root law comes from: information grows linearly with the sample, and error bars, being reciprocal square roots, shrink like $1/\sqrt{n}$. The probability part of this series built the score from densities; the probability chapter is where the expectation behind it is defined.
The score is also the natural object in optimisation. Near the peak of a log-likelihood you are doing gradient ascent on a random objective, and the argument that the noise in that gradient averages away is the argument that the estimate is consistent. The gradient-descent chapter sees the same score under a different name, and the Hessian it uses to set a step size is the curvature this part is about.
Information is curvature
Zoom the peak and read off the width
Differentiate the score once more and you get the curvature of the log-likelihood. Taking expectations and using the zero-mean property of the score gives the identity that ties the two readings together:
Read it as a shape statement. A large information means the log-likelihood curves downward hard as you leave the peak; a small information means it barely bends. The inverse information, $1/I(\theta)$, has units of squared parameter and is the natural error bar for one observation. The demo uses the simplest model whose information you can write down in closed form: a Bernoulli trial with success probability $p$. One draw has information
so $n$ draws carry $n\,I(p)$, and the MLE is just the observed success fraction $\hat p = k/n$. Drag the sample size and watch the log-likelihood, centred on $\hat p$ and drawn relative to its peak, with the dashed quadratic approximation that matches its curvature. The shaded strip is the interval $\hat p \pm 1/\sqrt{J(\hat p)}$, where $J$ is the observed information at the peak — the curvature actually measured on this sample, from the negative second derivative. At the peak of a Bernoulli likelihood that observed information is $n/(\hat p(1-\hat p))$, so $1/\sqrt{J} = \sqrt{\hat p(1-\hat p)/n}$, the familiar standard error. Changing $n$ reshapes the curve and rescales the axis; changing the true $p$ moves the dashed truth line.
Solid: log-likelihood of the sample, offset so its peak is zero. Dashed: the quadratic approximation $-\tfrac12 J(p)(p-\hat p)^2$. Shaded: $\hat p \pm 1/\sqrt{J}$.
The match is the point. Near its peak, any smooth log-likelihood looks quadratic, and the coefficient of that quadratic is the information. Estimating an error bar is therefore not a separate theory bolted onto estimation; it is reading the second derivative of the curve you already fitted. When people speak of "zooming into the peak", this is literally what they mean, and the zoom is a statement about second derivatives: the curvature of a surface is what tells you how fast it falls away from a maximum.
The Cramér–Rao bound
A floor under every unbiased estimator
Information is not just a description of the likelihood; it is a constraint. The Cramér–Rao lower bound states that for any unbiased estimator $\hat\theta$ of $\theta$,
The proof is a covariance calculation: the estimator and the score are both functions of the data, the estimator is unbiased, and the Cauchy–Schwarz inequality applied to their covariance yields the bound. What matters here is the consequence. The information in a sample puts a hard floor under the variance of any unbiased rule you could invent. Better estimators do not evade the bound; they approach it. An estimator that reaches it for every $\theta$ is called efficient, and its variance is exactly the reciprocal of the information.
The second demo draws this floor and the distribution sitting on top of it. It simulates many samples of size $n$, computes the MLE $\hat p = k/n$ for each, and builds the histogram of those estimates. Overlaid is the normal curve predicted by the central limit theorem, centred at $p$ with standard deviation $\sqrt{p(1-p)/n}$ — which is exactly the Cramér–Rao bound for the Bernoulli model. The dotted vertical lines mark one bound-width either side of the truth. As $n$ grows the histogram tightens onto the curve, and the curve tightens onto the truth at the rate $1/\sqrt{n}$.
Histogram of MLEs over repeated samples (solid bars), the CLT normal with the Cramér–Rao standard deviation (dashed curve), and the truth (dashed line).
The maximum-likelihood estimate is the estimator the bound was waiting for. It is asymptotically normal, asymptotically unbiased, and asymptotically efficient: its variance tends to $1/(n\,I(\theta))$ as $n$ grows. That is the precise sense in which "maximum likelihood is optimal" is true. For finite samples it can be biased — here $\hat p = k/n$ is exactly unbiased, but most interesting MLEs are not — and the bound applies to unbiased estimators, so a small-sample comparison needs care. The asymptotic claim is the one that survives, and it is the reason MLEs are the default in practice.
The bound is not a promise that your estimator is good, only a limit a bad one ignores; and it says nothing about biased estimators, some of which trade bias for variance and win in small samples. Its real use is diagnostic: an error bar smaller than $1/\sqrt{n\,I(\theta)}$ means something in the calculation is wrong.
Observed, expected, and design
Where the error bar comes from, and how to shrink it
Two information quantities appeared above and it is worth keeping them apart. The observed information is the negative second derivative of the log-likelihood at the estimate, computed from the data you have; it is what you report. The expected information is the average of that curvature over the distribution of the data at the true parameter; it is what the Cramér–Rao bound uses. They agree in expectation, and the demo shows the observed value tracking the plug-in estimate $\hat p(1-\hat p)/n$ as $n$ grows.
Information is additive over independent observations, and that single fact is the whole of experimental design: run two independent experiments and the information adds, so the variance floor halves when you double the data. But adding data is only one lever. Choose a design that makes each observation more informative — measure where the model is most sensitive, avoid settings where the score is nearly flat — and you raise $I(\theta)$ per sample instead. Reparametrisation matters too: information transforms by the square of a derivative, so the same data can look more or less informative depending on which parameter you report. That is the Jacobian rule again, and it is why a bound must always name its parameter — a theme developed in the numerics part of the linear algebra guide.
In multi-parameter problems the same story runs in matrix form. The information becomes a matrix, the negative Hessian of the log-likelihood; its inverse is the covariance floor; and the directions in which the information is small are the directions in which the experiment is ill-posed. Triangulating a 3D point from two noisy images is exactly this: the Fisher information of the ray geometry tells you which depths are pinned down and which are nearly unconstrained, and the triangulation chapter reads that matrix to decide whether a reconstruction is trustworthy. The same matrix, assembled over a trajectory, is what makes SLAM and bundle adjustment more than curve fitting: the SLAM chapter is the information matrix made into a map.
Where this shows up
Curvature as the universal error bar
Calibration as information
A robot calibrating its wheel radii and sensor offsets is estimating parameters from noisy motion, and the covariance of those estimates is the inverse Fisher information of the calibration trajectory. A path that excites every parameter makes the log-likelihood curve sharply in each direction; a path that only drives straight leaves the heading offset nearly flat, and the error bar for it explodes. The same point arrives from the drift side: poor conditioning is why dead reckoning diverges.
Conditioning is information
The inverse information is a covariance matrix, and the directions where it is large are the directions where the experiment, not the arithmetic, is weak. That is the same diagnosis as an ill-conditioned matrix in numerical linear algebra: a nearly flat log-likelihood is a nearly singular problem, and the error bar is the reciprocal of the curvature just as the condition number is the reciprocal of the smallest singular value.
Once you can read information off a curvature, a great many tools turn out to be the same tool. The Hessian that sets a Newton step in optimisation, the covariance a Kalman filter carries forward, the standard errors a regression printout reports, and the confidence interval around a benchmark score are all $1/\sqrt{\text{information}}$ in different clothing. The Jacobian and Hessian chapter supplies the calculus, and every place this site measures uncertainty it is, at bottom, measuring curvature. When you next see a standard error, ask what log-likelihood it came from; the answer is what makes it meaningful rather than decorative.
Further reading
Wasserman and Casella–Berger both state the Cramér–Rao theorem with its regularity conditions honestly; the information-geometry literature explains why the bound is the metric of a curved statistical space; and Cover and Thomas show the same quantity emerging from data-processing arguments. Read a chapter on optimal experimental design afterwards, and the vocabulary of D-optimality will look like matrix algebra rather than magic.
- Larry Wasserman, All of Statistics: A Concise Course in Statistical Inference, chapter 9 — the Cramér–Rao bound and efficiency.
- George Casella and Roger Berger, Statistical Inference, chapter 7 — information inequality, regularity conditions and the multiparameter case.
- Thomas Cover and Joy Thomas, Elements of Information Theory — why Fisher information is the local metric on a family of distributions.
- Shun-ichi Amari, Information Geometry and Its Applications — the log-likelihood as a surface and the bound as its curvature.
Cheat sheet
| Term | Meaning here |
|---|---|
| Score $s(\theta)$ | Derivative of the log-likelihood; mean zero, and the slope whose peak defines the MLE |
| Fisher information $I(\theta)$ | Variance of the score, equivalently the expected curvature of the log-likelihood |
| $n\,I(\theta)$ | Information in $n$ independent observations, because scores and their variances add |
| Observed information $J(\hat\theta)$ | Negative second derivative at the estimate, read from the data you have |
| Cramér–Rao bound | $\mathrm{Var}(\hat\theta) \ge 1/(n\,I(\theta))$ for any unbiased estimator |
| Efficiency | An estimator attains the bound, i.e. its variance equals $1/(n\,I(\theta))$ |
| MLE asymptotics | As $n\to\infty$ the MLE is normal with variance $1/(n\,I(\theta))$: efficient in the limit |
| Design | Raise $I$ per sample or add independent samples; information is the currency |