Covariance, correlation, and the ellipse
One variable is a cloud on a line; two variables are a cloud on a plane. This part puts two columns side by side and asks what a single number can honestly say about how they move together. You will drag a correlation and drag individual points and watch the covariance ellipse respond, and then meet four datasets that share every summary statistic yet look nothing alike. The point is not that summaries lie, but that a summary is a contract with a shape, and it pays to know which shapes it covers.
The question
What does it mean for two variables to move together?
A dataset with two columns is not two datasets. The rows pair an x with a y, so the object to study is a joint distribution over the plane: a rule that assigns probability to regions, not just to intervals on each axis. Squash that joint distribution onto the x axis and you get the marginal distribution of x; squash onto y and you get the marginal of y. The marginals are the two histograms you would draw if you threw away the pairing. Everything interesting is in what those marginals do not contain.
Two clouds can have identical x marginals and identical y marginals and still be completely different objects, because the pairing differs. Height and weight are positively associated: tall people tend to be heavier, so the cloud tilts up and to the right. The temperature outside and the number of layers of clothing are negatively associated. Height and shoe size are positively associated but with a lot of scatter. These are different joint distributions even though any single axis looks similar.
The first job of this part is to turn that tilt into a number. We want a summary of association that is positive when the two variables rise together, negative when one rises as the other falls, near zero when knowing one tells you little about the other, and measured on a scale that does not depend on the units we happened to choose. Two numbers do the work — covariance and its standardised cousin, correlation — and a third object, the covariance ellipse, makes them visible as a shape.
A cloud on a plane
Joint, marginal, and the pairing between them
The demo below draws points from a bivariate normal distribution. Both coordinates are centred at zero and have spread one, so the two marginals are identical standard normal histograms no matter what the correlation is. The only thing that changes as you move the slider is the pairing: how strongly a point's x predicts its y. At the left of the slider the cloud leans down and to the right, at the right it leans up and to the right, and at zero it is a circular blur with no preferred direction.
Each point is generated by taking two independent standard normal values and mixing them, so the association is real and not an artefact of the drawing order. You can also grab any individual point and drag it. A single dragged point is an outlier that the summary has to absorb: move it far enough and the correlation visibly sags or swings, which is a useful reminder that a correlation is computed from all the points and is not robust to a few of them moving.
The two marginal histograms are the shadows of the cloud on the axes. They are a genuine summary and a genuinely incomplete one: they carry the distribution of each variable separately and nothing whatsoever about how the two are paired. The joint distribution is the object; the marginals are two projections of it, and projections lose information. When a later part writes $f_{X,Y}(x,y)$ and then $f_X(x)$, the difference between those two symbols is exactly the difference between the cloud and its shadow.
Notice too that the correlation says nothing about the scale of the picture. Multiply every y by ten and the cloud becomes a tall thin blade, but the correlation is unchanged because it was standardised. Multiply a single point and the correlation does change, because standardisation is an average, not a per-point guarantee. That asymmetry between changing the units of a whole variable and changing one observation is the theme of the next step.
Covariance, correlation, and the ellipse
An average product, then a standardised version
Covariance is the average of a signed product. Centre each variable by subtracting its mean, multiply the two centred values at every point, and average. A point in the upper-right or lower-left quadrant has both centred values the same sign, so its product is positive; a point in the upper-left or lower-right has opposite signs, so its product is negative. The covariance is the balance of those contributions, which is why it is positive for an upward tilt and negative for a downward one.
The units give it away. If x is measured in metres and y in kilograms, the covariance is in metre-kilograms; change x to centimetres and the covariance is multiplied by a hundred while the relationship has not changed at all. A covariance of 4 is not "twice as strong" as a covariance of 2, because the number carries the units you chose. This is the one real defect of covariance as a measure of association, and correlation exists to fix it.
Correlation divides the covariance by the two standard deviations, which cancels the units and the arbitrary scale of each variable at once:
Because of the Cauchy–Schwarz inequality the result is trapped between $-1$ and $1$, and the endpoints are achieved only when the points lie exactly on a line, one rising with x and the other falling. A correlation of zero means the average positive and negative products cancel, which is a statement about linear association and only about linear association. Independence is stronger: if two variables are independent then their covariance is zero and their correlation is zero, but the converse fails. Take X symmetric about zero and let Y = X^2. Then Y is perfectly determined by X, yet the products pair each positive x with a positive y and each negative x with the same positive y, so the average product cancels and the covariance is zero. Zero covariance means "no straight-line trend", not "unrelated".
The covariance matrix collects both variances and the covariance into one symmetric array, and it is the object that PCA and least squares both consume:
Feed that matrix to a Cholesky factorisation and you get the lower-triangular $L$ with $\Sigma = LL^{T}$. The ellipse in the demo is the image of the unit circle under $L$, centred at the sample mean, scaled by a factor k. The k = 1 curve is the one-sigma contour of the fitted distribution and the dashed k = 2 curve is the two-sigma contour. Their tilt is the correlation made geometric, their widths are the two standard deviations, and the eigenvectors of $\Sigma$ point along the axes of the ellipse. That last fact is the seed of PCA: the direction of greatest spread is the leading eigenvector of the covariance matrix.
The same matrix answers a conditional question. Given the value of x, how much does y still vary? The residual variance is the Schur complement of the top-left entry of $\Sigma$,
which is also the residual variance of the least-squares line. The slope of that line is $\sigma_{XY}/\sigma_X^2$, and its intercept is chosen so the line passes through the mean point. Push the correlation toward $\pm 1$ and the ellipse collapses onto the line, the residual variance goes to zero, and prediction becomes exact. Push it to zero and the ellipse becomes an axis-aligned ellipse, the slope is zero, and knowing x tells you nothing about the expected y. The demo prints the covariance, the correlation, and that slope live, so you can check each sentence against the picture.
Drag any point. The solid ellipse is the one-sigma contour, the dashed one two-sigma, and the magenta line is the least-squares fit through the mean.
Identical summaries, different pictures
Anscombe's quartet
In 1973 the statistician Francis Anscombe published four small datasets with a devious property. Each has eleven points. Each has the same mean of x, the same mean of y, the same variance of x, the same variance of y, and the same correlation. Each even produces the same least-squares line, with the same intercept and slope. Every number a textbook asks you to compute is identical across the four. And yet the four datasets are four entirely different situations.
The first is a well-behaved linear blob. The second is a clean parabola, so the relationship is strong but not linear, and a correlation is simply the wrong summary. The third is a perfect line with a single outlier that drags the fit off the line by just the right amount to match the others. The fourth has one point with an extreme x that single-handedly supplies the entire slope; remove it and the remaining ten points have no relationship at all. The toggle below switches between them and holds the summary readout fixed while the picture changes completely.
Try to predict, before switching, which dataset the point belongs to. The exercise is the point: summary statistics are a compression, and compression is lossy. Covariance and correlation are exactly the right summaries when the joint distribution is roughly elliptical and the relationship is roughly linear, because in that case the ellipse really does contain the whole shape of the cloud. When the relationship curves, or when a couple of points carry the association, the numbers stay the same and stop meaning what you assumed. The honest workflow is to compute the numbers and look at the cloud, and the four-panel toggle is the reason why.
Anscombe's lesson has a modern cousin. The Datasaurus Dozen is a set of a dozen scatterplots, ranging from a dinosaur to a star, all engineered so that every dataset shares the same means, variances, and correlation to two decimal places. The construction is the same idea taken to a cartoonish extreme, and it exists to keep exactly this caution fresh in a world where it is easy to summarise a million points with a handful of decimals.
Four datasets, one identical set of summary statistics, one identical least-squares line.
Where this shows up
The covariance matrix is everywhere the plane is
Pose uncertainty as a covariance
A robot's estimate of where it is and which way it faces is never a point; it is a distribution, and the state of the art represents it with a covariance matrix. For a robot integrating wheel and inertial motion, forward motion and turning are uncertain and correlated, so the ellipse of possible poses tilts as the robot drives. Marginalising that matrix onto one coordinate asks "how unsure am I about position, ignoring heading?", which is a Schur complement, and the conditioning on a measurement is a Kalman update on the same object drawn here.
PCA is the ellipse, found
The low-rank PCA part takes the covariance matrix from this page and asks for its eigenvectors. Those eigenvectors are the axes of the ellipse: the long axis is the direction of greatest variance, the short axis is the direction of least. Projecting data onto the leading eigenvectors is the same as keeping only the widest directions of the cloud, which is why the covariance matrix rather than the raw data is the starting point for dimensionality reduction.
The same matrix is the bridge to regression. The least-squares slope printed by the demo is a covariance divided by a variance, and the residual variance is the Schur complement above; the least-squares part arrives at the identical formula from the normal equations, which is why the geometry of fitting a line and the geometry of a Gaussian ellipse are one subject seen from two sides. And when a model reports a scalar error, remember that the interesting structure is usually in the off-diagonal entries — the correlations — which the marginal uncertainties alone cannot express.
Further reading
The references below treat covariance as a geometric object and Anscombe as a mandatory caution rather than a footnote. If you take away one thing, take away the picture of an ellipse whose axes are the directions of greatest and least spread, and the memory of four clouds that no summary can tell apart.
- Francis Anscombe, "Graphs in Statistical Analysis", The American Statistician, 1973 — the original quartet and the argument for plotting first.
- Larry Wasserman, All of Statistics: A Concise Course in Statistical Inference, chapters 2 and 3 — joint, marginal and conditional distributions, covariance and correlation.
- Justin Matejka and George Fitzmaurice, "Same Stats, Different Graphs", 2017 — the Datasaurus Dozen, the modern version of Anscombe's point.
- 3Blue1Brown, the Essence of Statistics series — the visual style this guide follows for covariance and correlation.
Cheat sheet
| Term | Meaning here |
|---|---|
| Covariance | The average product of the two centred values, $\mathrm{Cov}(X,Y)=\mathbb{E}[(X-\mu_X)(Y-\mu_Y)]$; carries units |
| Correlation | The covariance divided by both standard deviations, $\rho=\mathrm{Cov}(X,Y)/(\sigma_X\sigma_Y)$; unitless, in $[-1,1]$ |
| Zero correlation | No linear trend. Independence implies it, but it does not imply independence |
| Joint / marginal | The distribution over the plane, and its two one-variable shadows on the axes |
| Covariance matrix | $\Sigma=\begin{bmatrix}\sigma_X^2 & \sigma_{XY}\\ \sigma_{XY} & \sigma_Y^2\end{bmatrix}$; variances on the diagonal, covariance off it |
| Covariance ellipse | The image of the unit circle under $L$ where $\Sigma=LL^{T}$; its axes are the eigenvectors of $\Sigma$ |
| Conditional variance | $\mathrm{Var}(Y\mid X)=\sigma_Y^2(1-\rho^2)$, the Schur complement and the residual variance of the fit |
| Anscombe's quartet | Four datasets with identical means, variances, correlation and fit, and four different shapes |