Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

The question

One cloud, one quadratic form

A random variable becomes a random vector the moment an experiment reports several numbers at once: a robot's pose is $(x,y,\theta)$, a small image patch is a grid of intensities, a fitted line is a slope and an intercept. The Gaussian is the default model for such a vector, for the same reason it is the default for one number — it is selected by the central limit theorem, it is the maximum-entropy choice once you fix a mean and a covariance, and — the practical reason — every operation you need stays inside the family.

For one variable the density is a bell centred at $\mu$ with width $\sigma$. The question of this part is what replaces the dome when the variable is a vector. The answer is a surface over the plane, a hill with a single summit at $\mu$, whose level sets are ellipses rather than circles. The ellipse shape is exactly the shape of the covariance, and the whole density can be reconstructed from just two quantities: where the summit is, and how the hill leans.

The reason the Gaussian is so tractable is the shape of the exponent. For a single variable the exponent is $-\tfrac12(x-\mu)^2/\sigma^2$, a quadratic in the deviation. In several dimensions the squared deviation becomes a squared length measured with the covariance matrix, $(x-\mu)^\top\Sigma^{-1}(x-\mu)$, a quadratic form. Quadratic forms are the one nonlinearity that linear algebra fully understands: their level sets are ellipses, their eigenvectors diagonalise them, and their inverse measures distance. That is why the multivariate Gaussian is a page of linear algebra wearing a probability hat.

Once you internalise that, the four workhorse operations fall out. Marginalising a Gaussian — forgetting a coordinate — just reads a diagonal block of $\Sigma$. Conditioning on a coordinate slices the hill and re-fits a Gaussian to the slice. Mahalanobis distance is the quadratic form itself, the natural ruler inside a cloud. Whitening multiplies by $\Sigma^{-1/2}$ and turns the tilted cloud into a round, standard one. We build each of these, drawn live.

💡 By the end of this part you'll see why a multivariate Gaussian is nothing but a mean and a symmetric positive-definite matrix, why that matrix's eigenvectors are the ellipse axes, and how marginals, conditionals, Mahalanobis distance and whitening are all the same quadratic form read in four different ways.
2

Density and quadratic form

The hill whose level sets are ellipses

Let $x$ be a vector in $\mathbb{R}^n$. A multivariate Gaussian is specified by a mean vector $\mu$ and a covariance matrix $\Sigma$, symmetric and positive definite. Its density is

$$ f(x) \;=\; (2\pi)^{-n/2}\,|\Sigma|^{-1/2}\,\exp\!\left(-\tfrac12\,(x-\mu)^\top \Sigma^{-1} (x-\mu)\right). $$

Read the formula in pieces. The prefactor is only a normalising constant — it is what makes the surface integrate to one, and it grows with the spread of the cloud, because a flatter hill has to cover the same unit of volume. All the shape lives in the exponent, and the exponent is $-\tfrac12$ times the quadratic form $Q(x) = (x-\mu)^\top\Sigma^{-1}(x-\mu)$. Because $\Sigma$ is positive definite, $\Sigma^{-1}$ is too, so $Q(x)\ge 0$ with equality only at the summit. The density is therefore a monotone function of a non-negative quadratic — a hill with a single peak.

Now ask what a level set looks like. Setting $Q(x) = k^2$ describes the set of points whose covariance-weighted squared distance from the centre is $k^2$. That is the equation of an ellipse centred at $\mu$ (an ellipsoid in $n$ dimensions), and its axes are the eigenvectors of $\Sigma$ with lengths proportional to the square roots of the eigenvalues. The constants $k=1,2,3$ are the familiar one-, two- and three-sigma contours: outside $k=1$ sits about $39\%$ of a two-dimensional Gaussian, outside $k=2$ about $13.5\%$, and outside $k=3$ about $1.1\%$.

Drag the three sliders below. The two variances set how far the cloud spreads along the coordinate axes; the correlation sets how strongly the two coordinates move together and therefore how much the ellipse leans. The right-hand panel is the same density as a surface: the quadratic form is literally the shape of the hill. Notice that changing the correlation never changes the peak height much — it slides probability around the saddle, it does not add or remove it.

Contours are level sets of the quadratic form, drawn from the eigenvalues and eigenvectors of $\Sigma$. Dots are samples from the density; the arrows are the eigenvectors scaled to the one-sigma semi-axes.

Plotly surface $z=f(x_1,x_2)$ for the same $\Sigma$. The ellipses of the panel above are the horizontal slices of this hill.

💡 The exponent is the geometry. Positive definiteness is exactly the statement that the hill has a single summit — if $\Sigma$ had a negative eigenvalue the quadratic form would dip and the "density" would not be integrable.
3

Σ as shape: eigen-decomposition

The spectral theorem for a covariance

Because $\Sigma$ is symmetric, the spectral theorem gives it an orthonormal eigenbasis. Stack the eigenvectors as columns of $V$ and put the eigenvalues on the diagonal of $\Lambda$:

$$ \Sigma \;=\; V\,\Lambda\,V^\top, \qquad V^\top V = I, \qquad \Lambda = \operatorname{diag}(\lambda_1,\dots,\lambda_n),\quad \lambda_i > 0. $$

This one line is the entire visual content of the covariance. The columns of $V$ are the directions the cloud leans along — the principal axes — and the $\lambda_i$ are the variances along those directions. Equivalently, the square roots $\sqrt{\lambda_i}$ are the semi-axis lengths of the one-sigma ellipse, and the axes point along the columns of $V$. The eigen-decomposition is not a computational curiosity; it is the canonical coordinate system in which the Gaussian becomes a product of independent one-dimensional Gaussians.

It also gives the cleanest way to build a Gaussian. Take a standard normal vector $z$ — independent components of variance one — and set $x = \mu + V\Lambda^{1/2}z$. Rotate a round cloud by $V$, stretch it by $\Lambda^{1/2}$, and translate by $\mu$. The covariance of the result is $V\Lambda^{1/2}\Lambda^{1/2}V^\top = \Sigma$. The matrix $A = V\Lambda^{1/2}$ is one square root of $\Sigma$; it is what the canvas below applies to the unit circle to draw the ellipse.

Two consequences are worth stating plainly. First, $\det\Sigma = \prod_i\lambda_i$, so the normalising constant is small exactly when the cloud is small; the density is high but the volume is low. Second, the directions of zero correlation are the eigen-directions, not the coordinate axes. A diagonal $\Sigma$ means the coordinate axes already are the eigenvectors; a tilted ellipse means they are not, and reading off $\sigma_{12}$ alone tells you a covariance but not a lean.

Move the sliders and watch the circle map to a leaning ellipse. The thick arrows are the images of the coordinate axes under the square root, i.e. the directions of the principal axes. The unit circle is the constant-Mahalanobis contour you would get if $\Sigma$ were the identity; the ellipse is what it becomes.

The unit circle mapped by $\Sigma^{1/2}=V\Lambda^{1/2}V^\top$ gives the one-sigma ellipse; the arrows are the mapped axes, which lie along the eigenvectors of $\Sigma$.

💡 Rotate, then stretch. Every Gaussian is a round standard cloud that has been rotated by the eigenvectors and stretched by the square roots of the eigenvalues.
4

Marginals and conditioning

Slicing the hill, and forgetting a coordinate

Two operations take a joint Gaussian apart. Marginalising ignores a coordinate; the result is still Gaussian, and its parameters are just the corresponding entries. If $x=(x_1,x_2)$ has mean $(\mu_1,\mu_2)$ and covariance with blocks $\Sigma_{11},\Sigma_{12},\Sigma_{21},\Sigma_{22}$, then

$$ x_1 \sim \mathcal{N}(\mu_1,\Sigma_{11}),\qquad x_2 \sim \mathcal{N}(\mu_2,\Sigma_{22}). $$

There is no correction term, no dependence on the other block: the marginal of a Gaussian is Gaussian, and the off-diagonal covariance $\Sigma_{12}$ disappears without a trace. That is a special property. For almost any other family, integrating out a correlated variable changes the shape of what remains; here the shape survives untouched. It is the same closure that makes the Gaussian the default in filtering: a measurement that only touches one coordinate leaves a Gaussian belief in that coordinate.

Conditioning is the opposite act: clamp a coordinate to a value $x_2=c$ and ask what the remaining coordinate looks like. For a Gaussian the answer is another Gaussian, with mean and covariance

$$ \mu_{1\mid 2} = \mu_1 + \Sigma_{12}\Sigma_{22}^{-1}(c-\mu_2), \qquad \Sigma_{1\mid 2} = \Sigma_{11} - \Sigma_{12}\Sigma_{22}^{-1}\Sigma_{21}. $$

Look hard at these two expressions. The conditional mean slides linearly with $c$: the more correlated the pair, the more a clamp on the second coordinate pulls the first. The conditional variance does not depend on $c$ at all — every horizontal slice through a Gaussian hill has the same width, only its centre moves. And that variance is the marginal variance with a subtraction: clamping removes the part of the uncertainty that the observed coordinate was explaining, so the conditional distribution is never wider than the marginal, and it is exactly as narrow as it would be if no correlation existed when $\Sigma_{12}=0$.

On the left, the 2D panel shows the slice line $x_2=c$; the dot is the conditional mean, whose position slides as $c$ moves. On the right, the 1D panel is the resulting distribution $p(x_1\mid x_2=c)$. Change the covariance sliders above and watch the slice's width respond — stronger correlation means a tighter conditional, because knowing $x_2$ now tells you a lot about $x_1$.

5

Mahalanobis and whitening

The right ruler, and the change of coordinates that proves it

How far is a point from the centre of a cloud? Euclidean distance is the wrong answer the moment the cloud is not round. A displacement along a long axis is cheap — many samples live out there — while the same displacement across a short axis is expensive. The correct ruler weights each direction by its variance, and the weighting is the quadratic form. The Mahalanobis distance is its square root:

$$ d_M(x) \;=\; \sqrt{(x-\mu)^\top \Sigma^{-1} (x-\mu)}. $$

So the level sets of constant Mahalanobis distance are exactly the ellipse contours of the density, and $f(x)$ is a monotone function of $d_M$. A point at the same Euclidean distance can sit inside or outside the one-sigma ellipse depending on which way it lies. Drag the probe on the canvas and watch the two outlines diverge: the dashed circle is Euclidean, the solid ellipse is Mahalanobis, and the readout reports both as well as the correct squared form. This is also the quantity a tracker minimises in gating and in the innovation term of a Kalman update — a measurement is accepted or rejected by how many sigmas away it is, not by how many metres.

There is a change of coordinates that makes the point undeniable. Because $\Sigma$ is symmetric and positive definite, it has a symmetric inverse square root $\Sigma^{-1/2} = V\Lambda^{-1/2}V^\top$. Apply it to the centred point:

$$ z \;=\; \Sigma^{-1/2}(x-\mu) \qquad\Longrightarrow\qquad z \sim \mathcal{N}(0, I). $$

This is whitening. It rotates the cloud by $V^\top$, scales each principal axis to unit variance, and rotates back — after which the cloud is perfectly round and its covariance is the identity. In whitened coordinates the Mahalanobis distance becomes ordinary Euclidean distance, $\|z\| = d_M(x)$, which is why whitening shows up wherever distances are compared: nearest-neighbour search, Gaussian-process kernels, and the decorrelation step inside many optimisers. The panel below shows the before and after. Same sample points, one linear map between them; the tilted ellipse becomes the unit circle and the probe's Mahalanobis distance becomes its radius.

Drag the probe. The dashed outline is the Euclidean circle of the probe's radius; the solid ellipse is the set of points at the probe's Mahalanobis distance.

Top: the correlated cloud with its one- and two-sigma ellipses. Bottom: the same samples after $z=\Sigma^{-1/2}(x-\mu)$; the cloud is round and its covariance is the identity.

6

Where this shows up

Ellipses under the machinery

Robotics / Vision

SLAM pose and landmark covariances

Every SLAM and bundle-adjustment system stores a Gaussian belief over poses and landmarks, and every update conditions it on a measurement using exactly the formulas of Step 4. The marginal blocks of $\Sigma$ are the uncertainty of the robot's trajectory; the conditional is the refined estimate after a loop closure.

Math

Symmetric matrices and the spectral theorem

The claim that the ellipse axes are orthogonal and real is the spectral theorem for symmetric matrices, applied to a positive-definite one. Covariance is the probability face of that theorem; the same diagonalisation drives PCA and the Hessian analysis of optimisation.

Math

Singular values of a data matrix

If $X$ is a matrix of centred samples, the eigenvectors of its covariance are the right singular vectors of $X$, and the eigenvalues are the squared singular values over $n$. That is the bridge from covariance to the SVD and to low-rank approximation.

Probability

Covariance as an object to read

The covariance part builds $\Sigma$ entry by entry and draws the one- and two-sigma ellipses from a draggable cloud. This part takes that matrix and reads it through its eigenvectors, conditionals and inverse — the same object, four more lenses.

7

Cheat sheet

Every formula in one place

IdeaFormulaReading
Density$(2\pi)^{-n/2}|\Sigma|^{-1/2}\exp(-\tfrac12(x-\mu)^\top\Sigma^{-1}(x-\mu))$A hill; the exponent is the quadratic form.
Quadratic form$Q(x)=(x-\mu)^\top\Sigma^{-1}(x-\mu)$Covariance-weighted squared distance from the centre.
Level sets$Q(x)=k^2$ is an ellipse$k=1,2,3$ are the sigma contours.
Eigen-decomposition$\Sigma=V\Lambda V^\top$, $V^\top V=I$Columns of $V$: axes. $\lambda_i$: variance along each.
Square root$\Sigma^{1/2}=V\Lambda^{1/2}V^\top$Maps the unit circle to the one-sigma ellipse.
Marginal$x_1\sim\mathcal{N}(\mu_1,\Sigma_{11})$Forget a coordinate: keep its diagonal block.
Conditional mean$\mu_{1\mid2}=\mu_1+\Sigma_{12}\Sigma_{22}^{-1}(c-\mu_2)$Slides linearly with the clamped value.
Conditional covariance$\Sigma_{1\mid2}=\Sigma_{11}-\Sigma_{12}\Sigma_{22}^{-1}\Sigma_{21}$Constant in $c$; marginal variance minus what was explained.
Mahalanobis distance$d_M=\sqrt{(x-\mu)^\top\Sigma^{-1}(x-\mu)}$The right ruler; level sets are the density ellipses.
Whitening$z=\Sigma^{-1/2}(x-\mu)\sim\mathcal{N}(0,I)$Rotate and scale the cloud to round and standard.
8

Further reading

Where to go deeper

9

Check your understanding

0/6 answered