Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

The question

Two numbers, one dance

Variance measures how far a single variable wanders from its mean. But most interesting quantities come in pairs. Height and weight, price and demand, two coordinates of a robot's pose, the activations of two neurons. Knowing that each one wanders is not enough; the important question is whether they wander in step. If you learn that someone is unusually tall, your prediction of their weight should move up too — and the strength of that should-be adjustment is exactly what covariance is measuring.

Here is the geometric picture before any symbols. Take every point in a cloud and slide the whole cloud so its centre of mass sits at the origin. Each point then has a signed horizontal displacement and a signed vertical displacement. Multiply those two displacements together. A point in the upper right has both positive, so the product is positive; so does a point in the lower left, because both are negative. Points in the other two quadrants have opposite signs and contribute negatively. Average the products across the cloud and you have a single number: positive when the cloud leans along the rising diagonal, negative when it leans the other way, near zero when it fills the axes symmetrically.

This single number has an obvious defect: it carries units. Measure height in centimetres rather than metres and the covariance grows by a hundred without the relationship changing at all. Divide it by the two standard deviations and the units cancel; that ratio is the correlation, a pure number trapped between −1 and 1. Finally, stack the two variances and the covariance into a 2×2 matrix and you have a compact object that remembers everything a Gaussian needs to know about the pair — including the orientation and lengths of the ellipse that encloses the cloud.

The rest of this part builds all three ideas in that order, one interactive panel at a time, and then turns to the trap. A correlation is only a summary, and four very different clouds can share the same summary exactly. Seeing that caveat is as important as seeing the formula.

💡 By the end of this part you'll see why covariance is the average product of centred deviations, why correlation is covariance in units of spread, how Σ packs a pair (and a whole cloud) into four numbers, and where a single correlation can lie about the shape it summarises.
2

Covariance

The average product of centred deviations

Let X and Y have means $\mu_X$ and $\mu_Y$. The deviation of a sample from its own mean is $X-\mu_X$. Covariance is the expected value of the product of the two deviations, and the algebra collapses it to a form that is usually easier to compute: expand the product, push the expectation through by linearity, and the two cross terms vanish because $E[X]=\mu_X$ and $E[Y]=\mu_Y$. What is left is the famous subtraction.

$$\operatorname{Cov}(X,Y)=\mathbb{E}\big[(X-\mu_X)(Y-\mu_Y)\big]=\mathbb{E}[XY]-\mathbb{E}[X]\,\mathbb{E}[Y].$$

Read the definition as a machine. Centre both variables, multiply them pointwise, average. Because multiplication is symmetric, $\operatorname{Cov}(X,Y)=\operatorname{Cov}(Y,X)$. Because expectation is linear, covariance is bilinear: pulling constants out of either slot just multiplies the result. And because the deviation of a constant is zero, adding a constant to either variable changes nothing at all — covariance is blind to where the cloud sits, only to how it leans.

That blindness is worth stating as a rule because it shows up constantly. Covariance is a property of the fluctuations, not the levels. Two thermometers that agree on the average but disagree on the fluctuations can have any covariance you like; two variables that differ only by a fixed offset have identical covariances with everything. And the special case $\operatorname{Cov}(X,X)$ is just the variance, which is why the whole covariance apparatus is best read as a generalisation of variance: it is the inner product of two centred variables, and variance is the inner product of a variable with itself.

The sign is the whole story in one glance. Positive covariance means the two variables tend to be above their means together and below them together; the cloud leans up and to the right. Negative covariance means one tends to be high while the other is low. Zero covariance — called uncorrelated — means the positive and negative products cancel, which can happen because the cloud is a symmetric blob, or for subtler reasons we will meet in the caveat below. Note carefully what covariance does not tell you: two variables can be perfectly dependent, with one an exact function of the other, yet have zero covariance, as long as the dependence is not monotone.

The panel below is the definition made drivable. A seeded cloud of ninety points is redrawn every time you move a slider. Drag the centre handle to translate the cloud: the covariance does not move, but the sample means do, and the ellipses slide along with the mean. Change $\rho$ to tilt the cloud, and $\sigma_X$, $\sigma_Y$ to stretch it, and watch the four numbers of $\Sigma$ respond.

The grey dots are the seeded sample. The blue and pink outlines are the 1σ and 2σ covariance ellipses. Drag the round centre handle to translate the cloud; the sliders tilt and stretch it.

Two habits are worth forming here. First, always centre before you interpret; a covariance computed around the origin rather than the mean picks up the offset and can flip sign for no good reason. Second, remember that covariance is an average of n products, so a single wild outlier can dominate it — which is precisely the lever the fourth cloud in the next demo pulls. The sample version below divides by n-1 rather than n for the familiar unbiasedness reason; in these clouds the distinction is invisible, but it matters as soon as samples are small.

3

Correlation

Covariance with the units taken out

Covariance grows with the spreads of both variables, which makes it impossible to compare across problems. The repair is to divide by those spreads. The result is the Pearson correlation, denoted $\rho$, a dimensionless number whose scale is fixed once and for all.

$$\rho_{XY}=\frac{\operatorname{Cov}(X,Y)}{\sigma_X\,\sigma_Y},\qquad -1\le\rho_{XY}\le 1.$$

That the correlation can never leave the interval [-1,1] is a restatement of the Cauchy–Schwarz inequality applied to the centred variables: the inner product of two vectors can never exceed the product of their lengths, and covariance is exactly the inner product of the centred X and Y vectors once the average is folded in. The two extremes are rigid. At $\rho=1$ the points lie exactly on a line of positive slope; at $\rho=-1$, exactly on a line of negative slope. Everything in between is a cloud whose thickness relative to its length is set by $\sqrt{1-\rho^2}$.

Correlation is invariant under a change of units, and more: if you rescale and shift either variable, aX+b and cY+d, the correlation is unchanged when a and c have the same sign and flips sign when they do not. This is what makes it comparable. A correlation of 0.8 between height in centimetres and weight in kilograms means the same thing as a correlation of 0.8 between the same variables measured in inches and pounds. The covariance, meanwhile, changed by the product of the conversion factors.

One warning bears repeating because it is the most common error in all of applied probability. Correlation measures linear association. Zero correlation does not mean independence. The classic counterexample is a variable and its own square on a symmetric range: Y=X^2 with X uniform on [-1,1] has $\operatorname{Cov}(X,Y)=0$ even though Y is a deterministic function of X. Independence is the stronger statement that the joint distribution factorises, which is the subject of Part 5; correlation only reads the second-moment shadow of that statement. The two coincide only in the special, and very useful, case of the joint Gaussian.

4

The covariance matrix

Four numbers that know the whole ellipse

Collect the variances on the diagonal and the covariance in the off-diagonal slots and you have the covariance matrix of the pair.

$$\Sigma=\begin{bmatrix}\sigma_X^2 & \operatorname{Cov}(X,Y)\\[2pt] \operatorname{Cov}(X,Y) & \sigma_Y^2\end{bmatrix}.$$

Because $\operatorname{Cov}(X,Y)=\operatorname{Cov}(Y,X)$, the matrix is symmetric — it equals its own transpose. That single fact unlocks the entire toolkit of symmetric matrices. A real symmetric matrix has real eigenvalues and an orthonormal basis of eigenvectors. Here those eigenvectors are the principal axes of the cloud and the eigenvalues are the variances along those axes. The level sets of the corresponding density are ellipses whose semi-axes point along the eigenvectors and have lengths proportional to the square roots of the eigenvalues. The interactive above draws exactly those ellipses at one and two standard deviations.

The matrix also answers the natural question about sums. The variance of X+Y is not the sum of the variances unless the covariance is zero; the cross term is the whole point.

$$\operatorname{Var}(X+Y)=\operatorname{Var}(X)+\operatorname{Var}(Y)+2\operatorname{Cov}(X,Y).$$

If the two variables move together, their sum is more variable than either alone, because the fluctuations reinforce. If they move oppositely, the sum is calmer — this cancellation is the entire idea behind diversification in a portfolio, and behind averaging correlated sensor readings. The matrix form of the same statement is that $\operatorname{Var}(X+Y)=[1\;\;1]\,\Sigma\,[1\;\;1]^{\mathsf T}$, a quadratic form that will reappear as the Mahalanobis distance in the next part.

There is one more property to carry forward. Because it is a covariance, $\Sigma$ is positive semidefinite: for any vector v, the number $v^{\mathsf T}\Sigma v$ is the variance of the linear combination v_X X + v_Y Y, and a variance cannot be negative. Equivalently, its determinant $\sigma_X^2\sigma_Y^2-\operatorname{Cov}(X,Y)^2$ is non-negative, which is the same Cauchy–Schwarz bound in matrix clothing. Negative determinant would mean a "spread" smaller than zero along some direction, and no joint distribution can do that.

Keep the geometry in view through all of this. A 2×2 symmetric positive semidefinite matrix is nothing more than a description of an ellipse: its eigenvalues give the squared lengths of the axes and its eigenvectors give the directions. Every operation on $\Sigma$ — adding, inverting, conditioning, taking square roots — is an operation on that ellipse. The next part makes the correspondence exact, but even here the picture pays for itself: a diagonal $\Sigma$ is an axis-aligned ellipse, an off-diagonal entry is a tilt, and a zero eigenvalue is an ellipse that has flattened to a line segment.

5

Four clouds, one correlation

A summary is not the shape

Correlation compresses a whole joint distribution into one number, and compression throws things away. The four panels below are Anscombe-style: each cloud has been transformed so that its sample correlation is the same, yet the joint shapes could hardly be more different. The first is a clean linear cloud with a little noise. The second curves upward — Y rises with X and then rises faster, a shape no straight line describes. The third is a tight little cluster with a single point flung far away; the outlier alone drags the correlation up. The fourth is bimodal, two clumps sitting on opposite sides with an empty gap between them.

All four report the same r. If you saw only the number you could not tell a genuine linear relationship from a parabola, a contaminated sample, or two populations mixed together. The lesson is not that correlation is useless — it is that you must always look at the cloud before you trust the number. This is the same discipline the earlier parts applied to a single variable: an average without a picture hides skew, fat tails and gaps. With two variables the temptation is stronger, because a single well-known number is easier to quote than a scatter plot.

There is a second, quieter lesson in the four panels. The straight and curved clouds were built from the same fifty standard normal draws; only the formula that turned X into Y changed. That is the sense in which a joint distribution is more than its two marginals: knowing the distribution of X and of Y separately leaves the coupling completely open, and it is $\Sigma$ — or, for non-Gaussian clouds, the full joint law — that supplies the missing information. For the Gaussian the covariance completes the picture; for these other shapes it only sketches it.

Linear

Curved

One outlier

Bimodal

The outlier panel deserves a second look, because it shows how fragile the second moment is. A handful of extreme points can move both the covariance and the variances, and since the correlation is a ratio of these, the effect does not simply vanish. Robust alternatives exist — Spearman's rank correlation replaces the values by their ranks and is unmoved by any monotone distortion, and Kendall's $\tau$ counts concordant and discordant pairs — but they too are single numbers, and they too can be gamed by shape. The honest conclusion is that no scalar can replace the scatter plot, and that correlation is best read alongside the variables' own spreads and a look at the data.

6

Where this shows up

Σ is the workhorse object

Three fields have independently converged on the same second-moment idea, which is a good sign it is the right abstraction.

Robotics

Estimating state, not just values

A filter on a robot's odometry tracks a mean state and a covariance matrix together. The diagonal says how uncertain each coordinate is; the off-diagonal says how errors are coupled, which decides how a single measurement should correct several coordinates at once. Drop the off-diagonals and the filter becomes overconfident in exactly the directions where the errors are correlated.

Linear algebra

Eigenvectors as principal axes

The covariance matrix is symmetric, so it diagonalises in an orthonormal basis. Its eigenvectors are the directions in which the cloud stretches and shrinks, and its eigenvalues rank those directions by variance. That is the entire geometric content of the next two links: compression by PCA is nothing but throwing away the small eigenvalues of Σ.

Optimisation

Curvature and correlated steps

The Hessian of a least-squares problem is a covariance-like matrix, and its correlation structure is why plain gradient descent zig-zags along narrow valleys. Preconditioning in nonlinear optimisation is the art of rescaling so that the effective covariance is close to the identity, at which point the elliptical valley becomes a round bowl.

Probability

When covariance tells the truth

Zero covariance and independence agree only for the joint Gaussian, where Σ captures the whole distribution. Everywhere else it is a partial description, which is why the multivariate Gaussian is the one place a covariance matrix is a complete summary — and why the next part is devoted to it.

7

Cheat sheet

Every formula in one place

IdeaFormulaReading
Covariance$\operatorname{Cov}(X,Y)=\mathbb{E}[(X-\mu_X)(Y-\mu_Y)]=\mathbb{E}[XY]-\mathbb{E}[X]\mathbb{E}[Y]$Average product of centred deviations; sign is the lean of the cloud.
Correlation$\rho_{XY}=\operatorname{Cov}(X,Y)/(\sigma_X\sigma_Y)$Covariance in units of spread; always in [-1,1].
Variance of a sum$\operatorname{Var}(X+Y)=\operatorname{Var}(X)+\operatorname{Var}(Y)+2\operatorname{Cov}(X,Y)$Correlated noise reinforces or cancels; independent noise adds.
Covariance matrix$\Sigma=\begin{bmatrix}\sigma_X^2 & \operatorname{Cov}(X,Y)\\ \operatorname{Cov}(X,Y) & \sigma_Y^2\end{bmatrix}$Symmetric, positive semidefinite; eigenvectors are the cloud's axes.
Linear maps$\operatorname{Cov}(aX+b,\;cY+d)=ac\,\operatorname{Cov}(X,Y)$Shifts do nothing; scaling multiplies by the product of the factors.
Ellipse axessemi-axis $k=\sqrt{\lambda_i}$, direction = eigenvectorThe 1σ and 2σ outlines drawn by $\operatorname{Prob.confEllipse}(\Sigma,\mu,k)$.
Uncorrelated$\operatorname{Cov}(X,Y)=0\iff\mathbb{E}[XY]=\mathbb{E}[X]\mathbb{E}[Y]$Weaker than independence; equal to it only for joint Gaussians.
Sample estimate$\widehat{\Sigma}_{xy}=\frac{1}{n-1}\sum_i (x_i-\bar{x})(y_i-\bar{y})$The n-1 denominator makes the estimate unbiased.
8

Further reading

Where to go deeper

9

Check your understanding

0/6 answered