Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

The question

How much does a sample know about a parameter?

Part 7 fitted a parameter by finding the top of a log-likelihood curve. The top tells you where the best guess is; it says nothing about how sharp that best guess is. A log-likelihood can be a narrow spike, in which case the data nearly determine the parameter, or a broad dome, in which case a wide range of values fits almost as well. The information in the sample is exactly that sharpness, and the whole of this part is one idea: the curvature of the log-likelihood at its peak measures how much the data know.

This is not a metaphor that happens to be useful. It is a theorem with a physical reading. Differentiate the log-likelihood once and you get the score, the direction and steepness of the slope. Differentiate again and you get the curvature, which tells you how quickly the slope falls away from the peak. The average squared score, and the average curvature, are the same number — the Fisher information — and that number sets a hard limit on how precisely any honest estimator can report the parameter.

Two questions organise the rest. First, what is the information, concretely, for a model simple enough to compute it by hand? Second, what does it forbid? The answer to the second is the Cramér–Rao lower bound: no unbiased estimator can beat the reciprocal of the information, and maximum likelihood, in the limit of large samples, exactly attains it. From there, experimental design becomes an information-accounting problem: to shrink an error bar you must add information, and adding data is only one of the ways to do it.

💡 By the end of this part you'll see why the second derivative of the log-likelihood is an error bar, why the score has zero mean and variance equal to that curvature, why the Cramér–Rao bound is a floor rather than a target, and why the maximum-likelihood estimate's sampling distribution lands on it as the sample grows.
2

The score function

The slope of the log-likelihood

Fix a model $p(x \mid \theta)$ with one unknown parameter $\theta$, and draw one observation $x$. The log-likelihood of that single observation is $\ell(\theta) = \log p(x \mid \theta)$, and its derivative

$$s(\theta) = \frac{\partial}{\partial\theta}\log p(x \mid \theta)$$

is the score. It measures how sensitively the log-probability of what you saw responds to a nudge of the parameter. At the maximum-likelihood estimate the score of the whole sample is zero — the peak is level — and the score of a single observation tells you which way that observation is pulling.

The score has a property that looks accidental and is not. Its expectation under the model is zero:

$$\mathbb{E}[\,s(\theta)\,] = \int \frac{\partial_\theta p(x\mid\theta)}{p(x\mid\theta)}\,p(x\mid\theta)\,dx = \partial_\theta \int p(x\mid\theta)\,dx = \partial_\theta 1 = 0.$$

The score is a random variable with mean zero, so all of its interesting content is in its variance. That variance is the Fisher information of a single observation,

$$I(\theta) = \mathrm{Var}\big[s(\theta)\big] = \mathbb{E}\big[s(\theta)^2\big],$$

and for $n$ independent observations the scores add, so their variances add too and the information is $n\,I(\theta)$. This is the first place the famous square-root law comes from: information grows linearly with the sample, and error bars, being reciprocal square roots, shrink like $1/\sqrt{n}$. The probability part of this series built the score from densities; the probability chapter is where the expectation behind it is defined.

The score is also the natural object in optimisation. Near the peak of a log-likelihood you are doing gradient ascent on a random objective, and the argument that the noise in that gradient averages away is the argument that the estimate is consistent. The gradient-descent chapter sees the same score under a different name, and the Hessian it uses to set a step size is the curvature this part is about.

3

Information is curvature

Zoom the peak and read off the width

Differentiate the score once more and you get the curvature of the log-likelihood. Taking expectations and using the zero-mean property of the score gives the identity that ties the two readings together:

$$I(\theta) = \mathbb{E}\big[s(\theta)^2\big] = -\,\mathbb{E}\!\left[\frac{\partial^2}{\partial\theta^2}\log p(x\mid\theta)\right].$$

Read it as a shape statement. A large information means the log-likelihood curves downward hard as you leave the peak; a small information means it barely bends. The inverse information, $1/I(\theta)$, has units of squared parameter and is the natural error bar for one observation. The demo uses the simplest model whose information you can write down in closed form: a Bernoulli trial with success probability $p$. One draw has information

$$I(p) = \frac{1}{p(1-p)},$$

so $n$ draws carry $n\,I(p)$, and the MLE is just the observed success fraction $\hat p = k/n$. Drag the sample size and watch the log-likelihood, centred on $\hat p$ and drawn relative to its peak, with the dashed quadratic approximation that matches its curvature. The shaded strip is the interval $\hat p \pm 1/\sqrt{J(\hat p)}$, where $J$ is the observed information at the peak — the curvature actually measured on this sample, from the negative second derivative. At the peak of a Bernoulli likelihood that observed information is $n/(\hat p(1-\hat p))$, so $1/\sqrt{J} = \sqrt{\hat p(1-\hat p)/n}$, the familiar standard error. Changing $n$ reshapes the curve and rescales the axis; changing the true $p$ moves the dashed truth line.

Solid: log-likelihood of the sample, offset so its peak is zero. Dashed: the quadratic approximation $-\tfrac12 J(p)(p-\hat p)^2$. Shaded: $\hat p \pm 1/\sqrt{J}$.

The match is the point. Near its peak, any smooth log-likelihood looks quadratic, and the coefficient of that quadratic is the information. Estimating an error bar is therefore not a separate theory bolted onto estimation; it is reading the second derivative of the curve you already fitted. When people speak of "zooming into the peak", this is literally what they mean, and the zoom is a statement about second derivatives: the curvature of a surface is what tells you how fast it falls away from a maximum.

⚠ Observed is not expected. $J(\hat p)$ is computed from the one sample in front of you; $n\,I(p)$ uses the true parameter. They are close at the MLE in large samples but differ for small $n$, and near $p=0$ or $p=1$ the observed information can blow up or mislead. The expected information is what the theory of the next section bounds; the observed information is what you actually have to report.
4

The Cramér–Rao bound

A floor under every unbiased estimator

Information is not just a description of the likelihood; it is a constraint. The Cramér–Rao lower bound states that for any unbiased estimator $\hat\theta$ of $\theta$,

$$\mathrm{Var}(\hat\theta) \;\ge\; \frac{1}{n\,I(\theta)}.$$

The proof is a covariance calculation: the estimator and the score are both functions of the data, the estimator is unbiased, and the Cauchy–Schwarz inequality applied to their covariance yields the bound. What matters here is the consequence. The information in a sample puts a hard floor under the variance of any unbiased rule you could invent. Better estimators do not evade the bound; they approach it. An estimator that reaches it for every $\theta$ is called efficient, and its variance is exactly the reciprocal of the information.

The second demo draws this floor and the distribution sitting on top of it. It simulates many samples of size $n$, computes the MLE $\hat p = k/n$ for each, and builds the histogram of those estimates. Overlaid is the normal curve predicted by the central limit theorem, centred at $p$ with standard deviation $\sqrt{p(1-p)/n}$ — which is exactly the Cramér–Rao bound for the Bernoulli model. The dotted vertical lines mark one bound-width either side of the truth. As $n$ grows the histogram tightens onto the curve, and the curve tightens onto the truth at the rate $1/\sqrt{n}$.

Histogram of MLEs over repeated samples (solid bars), the CLT normal with the Cramér–Rao standard deviation (dashed curve), and the truth (dashed line).

The maximum-likelihood estimate is the estimator the bound was waiting for. It is asymptotically normal, asymptotically unbiased, and asymptotically efficient: its variance tends to $1/(n\,I(\theta))$ as $n$ grows. That is the precise sense in which "maximum likelihood is optimal" is true. For finite samples it can be biased — here $\hat p = k/n$ is exactly unbiased, but most interesting MLEs are not — and the bound applies to unbiased estimators, so a small-sample comparison needs care. The asymptotic claim is the one that survives, and it is the reason MLEs are the default in practice.

The bound is not a promise that your estimator is good, only a limit a bad one ignores; and it says nothing about biased estimators, some of which trade bias for variance and win in small samples. Its real use is diagnostic: an error bar smaller than $1/\sqrt{n\,I(\theta)}$ means something in the calculation is wrong.

5

Observed, expected, and design

Where the error bar comes from, and how to shrink it

Two information quantities appeared above and it is worth keeping them apart. The observed information is the negative second derivative of the log-likelihood at the estimate, computed from the data you have; it is what you report. The expected information is the average of that curvature over the distribution of the data at the true parameter; it is what the Cramér–Rao bound uses. They agree in expectation, and the demo shows the observed value tracking the plug-in estimate $\hat p(1-\hat p)/n$ as $n$ grows.

Information is additive over independent observations, and that single fact is the whole of experimental design: run two independent experiments and the information adds, so the variance floor halves when you double the data. But adding data is only one lever. Choose a design that makes each observation more informative — measure where the model is most sensitive, avoid settings where the score is nearly flat — and you raise $I(\theta)$ per sample instead. Reparametrisation matters too: information transforms by the square of a derivative, so the same data can look more or less informative depending on which parameter you report. That is the Jacobian rule again, and it is why a bound must always name its parameter — a theme developed in the numerics part of the linear algebra guide.

In multi-parameter problems the same story runs in matrix form. The information becomes a matrix, the negative Hessian of the log-likelihood; its inverse is the covariance floor; and the directions in which the information is small are the directions in which the experiment is ill-posed. Triangulating a 3D point from two noisy images is exactly this: the Fisher information of the ray geometry tells you which depths are pinned down and which are nearly unconstrained, and the triangulation chapter reads that matrix to decide whether a reconstruction is trustworthy. The same matrix, assembled over a trajectory, is what makes SLAM and bundle adjustment more than curve fitting: the SLAM chapter is the information matrix made into a map.

6

Where this shows up

Curvature as the universal error bar

Robotics

Calibration as information

A robot calibrating its wheel radii and sensor offsets is estimating parameters from noisy motion, and the covariance of those estimates is the inverse Fisher information of the calibration trajectory. A path that excites every parameter makes the log-likelihood curve sharply in each direction; a path that only drives straight leaves the heading offset nearly flat, and the error bar for it explodes. The same point arrives from the drift side: poor conditioning is why dead reckoning diverges.

Numerics & ML

Conditioning is information

The inverse information is a covariance matrix, and the directions where it is large are the directions where the experiment, not the arithmetic, is weak. That is the same diagnosis as an ill-conditioned matrix in numerical linear algebra: a nearly flat log-likelihood is a nearly singular problem, and the error bar is the reciprocal of the curvature just as the condition number is the reciprocal of the smallest singular value.

Once you can read information off a curvature, a great many tools turn out to be the same tool. The Hessian that sets a Newton step in optimisation, the covariance a Kalman filter carries forward, the standard errors a regression printout reports, and the confidence interval around a benchmark score are all $1/\sqrt{\text{information}}$ in different clothing. The Jacobian and Hessian chapter supplies the calculus, and every place this site measures uncertainty it is, at bottom, measuring curvature. When you next see a standard error, ask what log-likelihood it came from; the answer is what makes it meaningful rather than decorative.

Further reading

Wasserman and Casella–Berger both state the Cramér–Rao theorem with its regularity conditions honestly; the information-geometry literature explains why the bound is the metric of a curved statistical space; and Cover and Thomas show the same quantity emerging from data-processing arguments. Read a chapter on optimal experimental design afterwards, and the vocabulary of D-optimality will look like matrix algebra rather than magic.

Cheat sheet

TermMeaning here
Score $s(\theta)$Derivative of the log-likelihood; mean zero, and the slope whose peak defines the MLE
Fisher information $I(\theta)$Variance of the score, equivalently the expected curvature of the log-likelihood
$n\,I(\theta)$Information in $n$ independent observations, because scores and their variances add
Observed information $J(\hat\theta)$Negative second derivative at the estimate, read from the data you have
Cramér–Rao bound$\mathrm{Var}(\hat\theta) \ge 1/(n\,I(\theta))$ for any unbiased estimator
EfficiencyAn estimator attains the bound, i.e. its variance equals $1/(n\,I(\theta))$
MLE asymptoticsAs $n\to\infty$ the MLE is normal with variance $1/(n\,I(\theta))$: efficient in the limit
DesignRaise $I$ per sample or add independent samples; information is the currency
7

Check your understanding

0/4 answered