Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

The question

What does it mean for two variables to move together?

A dataset with two columns is not two datasets. The rows pair an x with a y, so the object to study is a joint distribution over the plane: a rule that assigns probability to regions, not just to intervals on each axis. Squash that joint distribution onto the x axis and you get the marginal distribution of x; squash onto y and you get the marginal of y. The marginals are the two histograms you would draw if you threw away the pairing. Everything interesting is in what those marginals do not contain.

Two clouds can have identical x marginals and identical y marginals and still be completely different objects, because the pairing differs. Height and weight are positively associated: tall people tend to be heavier, so the cloud tilts up and to the right. The temperature outside and the number of layers of clothing are negatively associated. Height and shoe size are positively associated but with a lot of scatter. These are different joint distributions even though any single axis looks similar.

The first job of this part is to turn that tilt into a number. We want a summary of association that is positive when the two variables rise together, negative when one rises as the other falls, near zero when knowing one tells you little about the other, and measured on a scale that does not depend on the units we happened to choose. Two numbers do the work — covariance and its standardised cousin, correlation — and a third object, the covariance ellipse, makes them visible as a shape.

💡 By the end of this part you'll see why covariance is an average product of deviations and therefore depends on units, why correlation divides that dependence away and lands in $[-1,1]$, why the covariance matrix draws an ellipse that is the unit circle pushed through the data's own stretching, and why a summary can be exactly right and still describe four very different pictures.
2

A cloud on a plane

Joint, marginal, and the pairing between them

The demo below draws points from a bivariate normal distribution. Both coordinates are centred at zero and have spread one, so the two marginals are identical standard normal histograms no matter what the correlation is. The only thing that changes as you move the slider is the pairing: how strongly a point's x predicts its y. At the left of the slider the cloud leans down and to the right, at the right it leans up and to the right, and at zero it is a circular blur with no preferred direction.

Each point is generated by taking two independent standard normal values and mixing them, so the association is real and not an artefact of the drawing order. You can also grab any individual point and drag it. A single dragged point is an outlier that the summary has to absorb: move it far enough and the correlation visibly sags or swings, which is a useful reminder that a correlation is computed from all the points and is not robust to a few of them moving.

The two marginal histograms are the shadows of the cloud on the axes. They are a genuine summary and a genuinely incomplete one: they carry the distribution of each variable separately and nothing whatsoever about how the two are paired. The joint distribution is the object; the marginals are two projections of it, and projections lose information. When a later part writes $f_{X,Y}(x,y)$ and then $f_X(x)$, the difference between those two symbols is exactly the difference between the cloud and its shadow.

Notice too that the correlation says nothing about the scale of the picture. Multiply every y by ten and the cloud becomes a tall thin blade, but the correlation is unchanged because it was standardised. Multiply a single point and the correlation does change, because standardisation is an average, not a per-point guarantee. That asymmetry between changing the units of a whole variable and changing one observation is the theme of the next step.

3

Covariance, correlation, and the ellipse

An average product, then a standardised version

Covariance is the average of a signed product. Centre each variable by subtracting its mean, multiply the two centred values at every point, and average. A point in the upper-right or lower-left quadrant has both centred values the same sign, so its product is positive; a point in the upper-left or lower-right has opposite signs, so its product is negative. The covariance is the balance of those contributions, which is why it is positive for an upward tilt and negative for a downward one.

$$\mathrm{Cov}(X,Y) = \mathbb{E}\big[(X-\mu_X)(Y-\mu_Y)\big] \;\approx\; \frac{1}{n-1}\sum_{i=1}^{n}(x_i-\bar{x})(y_i-\bar{y})$$

The units give it away. If x is measured in metres and y in kilograms, the covariance is in metre-kilograms; change x to centimetres and the covariance is multiplied by a hundred while the relationship has not changed at all. A covariance of 4 is not "twice as strong" as a covariance of 2, because the number carries the units you chose. This is the one real defect of covariance as a measure of association, and correlation exists to fix it.

Correlation divides the covariance by the two standard deviations, which cancels the units and the arbitrary scale of each variable at once:

$$\rho = \mathrm{Corr}(X,Y) = \frac{\mathrm{Cov}(X,Y)}{\sigma_X\,\sigma_Y}$$

Because of the Cauchy–Schwarz inequality the result is trapped between $-1$ and $1$, and the endpoints are achieved only when the points lie exactly on a line, one rising with x and the other falling. A correlation of zero means the average positive and negative products cancel, which is a statement about linear association and only about linear association. Independence is stronger: if two variables are independent then their covariance is zero and their correlation is zero, but the converse fails. Take X symmetric about zero and let Y = X^2. Then Y is perfectly determined by X, yet the products pair each positive x with a positive y and each negative x with the same positive y, so the average product cancels and the covariance is zero. Zero covariance means "no straight-line trend", not "unrelated".

The covariance matrix collects both variances and the covariance into one symmetric array, and it is the object that PCA and least squares both consume:

$$\Sigma = \begin{bmatrix} \sigma_X^2 & \sigma_{XY} \\ \sigma_{XY} & \sigma_Y^2 \end{bmatrix} = \begin{bmatrix} \sigma_X^2 & \rho\,\sigma_X\sigma_Y \\ \rho\,\sigma_X\sigma_Y & \sigma_Y^2 \end{bmatrix}$$

Feed that matrix to a Cholesky factorisation and you get the lower-triangular $L$ with $\Sigma = LL^{T}$. The ellipse in the demo is the image of the unit circle under $L$, centred at the sample mean, scaled by a factor k. The k = 1 curve is the one-sigma contour of the fitted distribution and the dashed k = 2 curve is the two-sigma contour. Their tilt is the correlation made geometric, their widths are the two standard deviations, and the eigenvectors of $\Sigma$ point along the axes of the ellipse. That last fact is the seed of PCA: the direction of greatest spread is the leading eigenvector of the covariance matrix.

The same matrix answers a conditional question. Given the value of x, how much does y still vary? The residual variance is the Schur complement of the top-left entry of $\Sigma$,

$$\mathrm{Var}(Y \mid X) = \sigma_Y^2 - \frac{\sigma_{XY}^2}{\sigma_X^2} = \sigma_Y^2\,(1-\rho^2),$$

which is also the residual variance of the least-squares line. The slope of that line is $\sigma_{XY}/\sigma_X^2$, and its intercept is chosen so the line passes through the mean point. Push the correlation toward $\pm 1$ and the ellipse collapses onto the line, the residual variance goes to zero, and prediction becomes exact. Push it to zero and the ellipse becomes an axis-aligned ellipse, the slope is zero, and knowing x tells you nothing about the expected y. The demo prints the covariance, the correlation, and that slope live, so you can check each sentence against the picture.

Drag any point. The solid ellipse is the one-sigma contour, the dashed one two-sigma, and the magenta line is the least-squares fit through the mean.

one-sigma ellipse two-sigma ellipse least-squares line
The slider regenerates the cloud from a fixed pair of standard-normal draws. Dragging edits individual points in place.
⚠ Correlation is not causation, and it is also not curvature. A correlation near zero can hide a strong curved relationship, and a correlation near one tells you the points are close to a line, not that one variable drives the other. Before trusting either number, plot the cloud — which is precisely the lesson the next step makes unforgettable.
4

Identical summaries, different pictures

Anscombe's quartet

In 1973 the statistician Francis Anscombe published four small datasets with a devious property. Each has eleven points. Each has the same mean of x, the same mean of y, the same variance of x, the same variance of y, and the same correlation. Each even produces the same least-squares line, with the same intercept and slope. Every number a textbook asks you to compute is identical across the four. And yet the four datasets are four entirely different situations.

The first is a well-behaved linear blob. The second is a clean parabola, so the relationship is strong but not linear, and a correlation is simply the wrong summary. The third is a perfect line with a single outlier that drags the fit off the line by just the right amount to match the others. The fourth has one point with an extreme x that single-handedly supplies the entire slope; remove it and the remaining ten points have no relationship at all. The toggle below switches between them and holds the summary readout fixed while the picture changes completely.

Try to predict, before switching, which dataset the point belongs to. The exercise is the point: summary statistics are a compression, and compression is lossy. Covariance and correlation are exactly the right summaries when the joint distribution is roughly elliptical and the relationship is roughly linear, because in that case the ellipse really does contain the whole shape of the cloud. When the relationship curves, or when a couple of points carry the association, the numbers stay the same and stop meaning what you assumed. The honest workflow is to compute the numbers and look at the cloud, and the four-panel toggle is the reason why.

Anscombe's lesson has a modern cousin. The Datasaurus Dozen is a set of a dozen scatterplots, ranging from a dinosaur to a star, all engineered so that every dataset shares the same means, variances, and correlation to two decimal places. The construction is the same idea taken to a cartoonish extreme, and it exists to keep exactly this caution fresh in a world where it is easy to summarise a million points with a handful of decimals.

Four datasets, one identical set of summary statistics, one identical least-squares line.

The readout does not change when you switch. Only the cloud does.
5

Where this shows up

The covariance matrix is everywhere the plane is

Robotics

Pose uncertainty as a covariance

A robot's estimate of where it is and which way it faces is never a point; it is a distribution, and the state of the art represents it with a covariance matrix. For a robot integrating wheel and inertial motion, forward motion and turning are uncertain and correlated, so the ellipse of possible poses tilts as the robot drives. Marginalising that matrix onto one coordinate asks "how unsure am I about position, ignoring heading?", which is a Schur complement, and the conditioning on a measurement is a Kalman update on the same object drawn here.

ML / AI

PCA is the ellipse, found

The low-rank PCA part takes the covariance matrix from this page and asks for its eigenvectors. Those eigenvectors are the axes of the ellipse: the long axis is the direction of greatest variance, the short axis is the direction of least. Projecting data onto the leading eigenvectors is the same as keeping only the widest directions of the cloud, which is why the covariance matrix rather than the raw data is the starting point for dimensionality reduction.

The same matrix is the bridge to regression. The least-squares slope printed by the demo is a covariance divided by a variance, and the residual variance is the Schur complement above; the least-squares part arrives at the identical formula from the normal equations, which is why the geometry of fitting a line and the geometry of a Gaussian ellipse are one subject seen from two sides. And when a model reports a scalar error, remember that the interesting structure is usually in the off-diagonal entries — the correlations — which the marginal uncertainties alone cannot express.

Further reading

The references below treat covariance as a geometric object and Anscombe as a mandatory caution rather than a footnote. If you take away one thing, take away the picture of an ellipse whose axes are the directions of greatest and least spread, and the memory of four clouds that no summary can tell apart.

Cheat sheet

TermMeaning here
CovarianceThe average product of the two centred values, $\mathrm{Cov}(X,Y)=\mathbb{E}[(X-\mu_X)(Y-\mu_Y)]$; carries units
CorrelationThe covariance divided by both standard deviations, $\rho=\mathrm{Cov}(X,Y)/(\sigma_X\sigma_Y)$; unitless, in $[-1,1]$
Zero correlationNo linear trend. Independence implies it, but it does not imply independence
Joint / marginalThe distribution over the plane, and its two one-variable shadows on the axes
Covariance matrix$\Sigma=\begin{bmatrix}\sigma_X^2 & \sigma_{XY}\\ \sigma_{XY} & \sigma_Y^2\end{bmatrix}$; variances on the diagonal, covariance off it
Covariance ellipseThe image of the unit circle under $L$ where $\Sigma=LL^{T}$; its axes are the eigenvectors of $\Sigma$
Conditional variance$\mathrm{Var}(Y\mid X)=\sigma_Y^2(1-\rho^2)$, the Schur complement and the residual variance of the fit
Anscombe's quartetFour datasets with identical means, variances, correlation and fit, and four different shapes
6

Check your understanding

0/4 answered