Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

The question

What does a bell curve become when it has more than one dimension?

In one dimension a Gaussian is a bell curve: a location $\mu$ and a spread $\sigma^2$ are enough to draw it, to compute the probability of any interval, and to answer every question you can ask. Move to $n$ dimensions and the object becomes a density over $\mathbb{R}^n$ with two parameters, a mean vector $\mu$ and a symmetric positive definite matrix $\Sigma$. Everything you already trust about the one-dimensional Gaussian survives; what is new is that the variables can now be correlated, and the extra numbers that record those correlations are the off-diagonal entries of $\Sigma$. Those entries turn a circular hill into a tilted ridge, and almost every fact in this part is a consequence of the tilt.

The multivariate normal, or MVN, is the single most reused distribution in applied mathematics. States, sensor errors, filter innovations, model parameters and latent codes are all routinely assumed to be jointly Gaussian, not because the world is truly Gaussian but because the family is closed under the operations we perform on it. Add two Gaussians and you get a Gaussian. Marginalise a Gaussian and you get a Gaussian. Condition one coordinate on another and you get a Gaussian. Apply an affine map and you still have a Gaussian. That closure is the workhorse, and the purpose of this part is to make each of those four sentences something you can see.

The plan is geometric. Section 2 writes down the joint density and identifies its exponent as a quadratic form, so that the level sets are ellipsoids. Section 3 hands you the covariance ellipse and lets you drag it. Sections 4 and 5 read the two marginals and the conditional distribution off the same ellipse, and name the Schur complement that governs preconditioning. Section 6 applies an affine map and confirms that Gaussianity is preserved, and shows the standard way to turn independent standard normals into a sample from any MVN.

💡 By the end of this part you'll see why the joint density is a quadratic form in the exponent, why the covariance matrix draws an ellipse whose axes are its eigenvectors, why the marginals are Gaussian, why conditioning on one coordinate leaves a Gaussian whose mean is linear in the conditioning value and whose variance does not depend on it, and why affine maps and sampling both reduce to the same Cholesky factor.
2

The joint density and the quadratic form

One exponent, one ellipse

Write a point as a column vector $x = (x_1, \dots, x_n)^{\mathsf T}$. The multivariate normal density with mean $\mu$ and covariance $\Sigma$ is

$$f(x) \;=\; \frac{1}{(2\pi)^{n/2}\,\sqrt{\det \Sigma}}\;\exp\!\Big(-\tfrac12\,(x-\mu)^{\mathsf T}\Sigma^{-1}(x-\mu)\Big).$$

Read it in three pieces. The front factor is nothing but the normalising constant: the determinant of $\Sigma$ measures the $n$-dimensional volume the distribution fills, and dividing by its square root makes the density integrate to one, exactly as $1/(\sigma\sqrt{2\pi})$ does in one dimension. The exponent is a quadratic form in the centred variable, measured through the inverse covariance. And the symmetry of $\Sigma$ is what guarantees that form is a genuine squared length, so it can never go negative. The combination $(x-\mu)^{\mathsf T}\Sigma^{-1}(x-\mu)$ is the squared Mahalanobis distance from $x$ to the mean, and it is the natural notion of distance for this distribution.

Because the density depends on $x$ only through that quadratic form, its level sets are the sets where the form takes a fixed value. Those sets are ellipsoids centred at $\mu$: in the plane, ellipses. The special one where the quadratic form equals one is the one-sigma contour, and the one where it equals four is the two-sigma contour. The covariance matrix does not merely describe the spread; it is the shape of those ellipses, with the inverse appearing because a narrow direction means a large penalty for moving along it. The same quadratic form is the Hessian of the exponent, scaled by $\Sigma^{-1}$, which is why the curvature of the log-density and the covariance are reciprocal objects — the connection developed alongside the Jacobian and Hessian chapter of the calculus guide.

For $n=1$ the inverse covariance is $1/\sigma^2$ and the ellipse degenerates to an interval $[\mu-\sigma, \mu+\sigma]$. For $n=2$, which is where the rest of this part lives, the ellipse has a size, an aspect ratio and an orientation, and all three are encoded in the entries of a two-by-two matrix. The demo in the next section puts those entries under a pair of draggable handles, so the geometry is something you manipulate rather than infer.

3

Covariance as an ellipse

Drag the matrix, watch the density

A symmetric positive definite two-by-two matrix can always be written as $\Sigma = LL^{\mathsf T}$ with $L$ lower triangular, which is the Cholesky factorisation. The columns of $L$ are the vectors that carry the standard unit circle, applied through $L$, onto the one-sigma ellipse: the ellipse is literally the image of the circle under $L$. That is why the demos in this guide draw covariance ellipses by feeding the Cholesky factor to the circle construction. The diagonal entries of $\Sigma$ are the variances $\sigma_1^2$ and $\sigma_2^2$; the shared off-diagonal entry is the covariance. Divide it by the two standard deviations and you have the correlation $\rho$, a unitless number in $[-1,1]$ that says how strongly the two coordinates lean together.

The demo below starts with a positively correlated Gaussian. The two round handles at the ends of the coloured axes are the columns of $L$, so dragging one stretches the ellipse along that direction and dragging the other rotates and tilts it; the covariance matrix, the correlation and the two standard deviations update live. Sliders do the same job in the more familiar coordinates: turn $\rho$ up and the ellipse narrows and tilts toward the rising diagonal; push it to zero and the ellipse becomes axis-aligned; make one standard deviation large and the ellipse is a long blade along the corresponding axis.

Solid ellipse: one-sigma contour. Dashed: two-sigma. The dashed line is the conditional mean; the magenta line is the slice $x_1 = a$. Marginals sit on the top and right margins.

Drag the two covariant handles to set the covariance directly, or move the sliders. The magenta handle drags the slice.

The ellipse is also a picture of conditioning to come. The eigenvectors of $\Sigma$ point along its axes, so the longest axis is the direction in which the distribution is most uncertain and the shortest is the direction in which it is most concentrated. That is the same statement PCA makes, one part of this site away in the low-rank PCA chapter. Near-singular covariances make the ellipse needle-thin, and the Cholesky factor becomes a numerically delicate object to compute — the practical worry treated in the numerics chapter and the reason filters carry a condition number.

Plotly contour of the same joint density, with 220 draws from the matching MVN overlaid. Drag or slide and both views stay in step.

4

Marginals on the margins

Shadows on the axes

Suppose you care only about $x_1$ and are willing to ignore $x_2$ entirely. To get its distribution you integrate the joint density over every possible value of $x_2$, which is what "marginalising" means. For the MVN this integration can be done in closed form, and the result is exactly what you would hope: the marginal of $x_1$ is a one-dimensional Gaussian with mean $\mu_1$ and variance $\Sigma_{11}$, the top-left entry of the covariance matrix. The same statement holds for any subset of coordinates — the marginal over a subset is again Gaussian, with the corresponding sub-block of $\Sigma$ as its covariance. In the demo the marginal of $x_1$ is drawn along the top margin as the light curve, and the marginal of $x_2$ along the right margin. They are the shadows the ellipse casts on the two axes.

Marginalising is easy here for a structural reason. The quadratic form in the exponent can be rewritten so that the variable being integrated out appears only in a square, and a Gaussian integral over a square is a constant we already know how to compute. Nothing about the shape of the density is disturbed; the remaining variable keeps its Gaussian form. This is the first of the four closure properties, and it is what makes Gaussian models tractable: you can throw away coordinates you do not need without leaving the family.

What marginalising destroys is the pairing. The two marginals tell you everything about each coordinate separately and nothing whatsoever about how they move together, so two very different joint distributions — one tightly correlated, one independent — can have identical marginals. In the demo, drag $\rho$ to either extreme with the standard deviations fixed: the top and side curves do not budge, while the ellipse and the contours rotate beneath them. The correlation lives entirely in the off-diagonal part of $\Sigma$, which the marginals do not see. The companion guide's covariance-and-correlation part makes the same point with a scatter cloud; here it is made with a density.

5

Conditionals and the Schur complement

Slicing the ellipse

Now suppose you learn that $x_1$ is exactly $a$ — a sensor reading, a fixed value, a design point. The distribution of $x_2$ given this information is the conditional distribution, obtained by cutting the joint density along the vertical line $x_1 = a$ and renormalising what is left. Because a Gaussian's exponent is quadratic, the sliced cross-section is again a quadratic in $x_2$, so the conditional distribution is Gaussian. Its two parameters are

$$\mu_{2\mid 1} = \mu_2 + \Sigma_{21}\Sigma_{11}^{-1}(a-\mu_1), \qquad \Sigma_{2\mid 1} = \Sigma_{22} - \Sigma_{21}\Sigma_{11}^{-1}\Sigma_{12}.$$

The conditional mean is linear in the conditioning value: as $a$ moves, the centre of the sliced bell slides along a straight line, which is precisely the regression line of $x_2$ on $x_1$. That line is drawn in the demo as the dashed diagonal, and the magenta marker shows where the slice currently sits. The coefficient $\Sigma_{21}\Sigma_{11}^{-1}$ is the least-squares slope, which is why conditioning a Gaussian and fitting a line by least squares arrive at the same formula — a correspondence the least-squares part derives from the normal equations.

The conditional variance is the surprise, and it is the cleaner half of the result. It does not depend on $a$ at all. Whatever value you condition on, the slice leaves exactly the same residual spread, given by the Schur complement of the top-left block inside $\Sigma$. In the two-by-two case it simplifies to $\Sigma_{2\mid 1} = \sigma_2^2(1-\rho^2)$, the residual variance after the best linear prediction is removed. Push $|\rho|$ toward one and it collapses to zero: a perfect line, exact prediction. Push $\rho$ to zero and it equals the full marginal variance $\sigma_2^2$: knowing $x_1$ tells you nothing. The demo draws the marginal of $x_2$ and the conditional on the same right margin so the difference is visible — the conditional is the narrower, shifted curve.

That Schur complement is the geometric heart of the matter. Slicing the ellipse at $x_1=a$ always produces the same vertical extent, no matter where the slice is placed, and that fixed extent is $\Sigma_{2\mid 1}$. The same object reappears whenever a large covariance matrix is split into a block you know and a block you want to predict; it is the algebraic engine behind Gaussian elimination, behind the Kalman filter's conditioning step, and behind the marginalisation of a joint distribution onto a subset of its coordinates. Seeing it as "the width of every slice of this ellipse" makes the formula memorable rather than something to look up.

⚠ Conditional is not marginal. The marginal of $x_2$ averages over all values of $x_1$ and is the wider curve; the conditional fixes $x_1$ and is the narrower one. A common mistake is to read a marginal interval as if it were a conditional one, which understates how tightly a specific slice is pinned down — or overstates it when $\rho$ is near zero, where the two curves coincide.
6

Affine maps and sampling

Gaussianity in, Gaussianity out

The last closure property is the one that makes the family so useful in engineering. Take a Gaussian vector $x \sim \mathcal{N}(\mu, \Sigma)$ and form the affine image $y = Ax + b$. The result is again Gaussian, with mean $A\mu + b$ and covariance $A\Sigma A^{\mathsf T}$. A linear combination of jointly Gaussian variables is Gaussian, full stop; no approximation is involved. The demo below starts from a standard normal cloud $z \sim \mathcal{N}(0, I)$, drawn faintly, and pushes it through a matrix $A$ you can edit or replace with a preset. The transformed cloud is coloured, and the ellipse is the image of the unit circle under $A$. Rotate, shear or scale $A$ and the cloud and the ellipse transform together: the Gaussian shape is never lost, only stretched and turned.

This is exactly how you sample from an MVN. If $z$ is a vector of independent standard normals and $\Sigma = LL^{\mathsf T}$, then setting $x = \mu + Lz$ produces a draw with mean $\mu$ and covariance $\Sigma$, because $L z$ has covariance $L\,I\,L^{\mathsf T} = \Sigma$. Cholesky is the bridge: it is the unique lower-triangular square root that turns a round cloud into a tilted one. The function Stats.sample.mvnormal used on this page does precisely this internally — draw standard normals, multiply by the Cholesky factor, add the mean — and the coloured points in the contour view are draws from it.

Faint: $z \sim \mathcal{N}(0, I)$ with the unit circle. Coloured: $x = Az$ with the image ellipse $\Sigma = AA^{\mathsf T}$. A Gaussian stays a Gaussian.

Two consequences are worth carrying away. First, the mean shifts affinely and the covariance transforms by $A\Sigma A^{\mathsf T}$, so a Gaussian is completely described by how its first two moments move — there are no higher moments to track because they are all determined. Second, sampling and affine transformation are the same operation run in opposite directions, which is why a single Cholesky factor serves both the "generate a draw" and the "propagate uncertainty" tasks. The probability chapter of the calculus guide, probability in motion, treats the density-transformation view of the same fact.

7

Where this shows up

One distribution, an entire engineering stack

Robotics & Vision

Gaussian state estimation

A robot's belief about its pose is a Gaussian, and every measurement is a slice. SLAM systems keep the joint distribution of the trajectory and the map as a large covariance matrix and update it by conditioning; the SLAM chapter is this page with thousands of coordinates instead of two. The four closure properties you just watched are exactly what lets those systems keep a single belief rather than a pile of samples.

ML / AI

The covariance matrix and PCA

The ellipse on this page is the covariance matrix made visible, and asking for its axes is asking for its eigenvectors. The low-rank PCA chapter takes the same matrix and keeps only the widest directions, which is the ellipse's longest axis promoted to a summary of the whole cloud. Dimensionality reduction, whitening and latent-variable models all begin from this object.

The Schur complement reappears the moment a filter runs. Predicting forward moves the mean and covariance through an affine map; correcting with a measurement conditions on a slice; and the filter's covariance after the update is a Schur complement of the joint distribution of state and measurement. The next parts turn each of those steps into its own demo — the recursive belief update, then the linear and extended Kalman filters that assume Gaussianity outright, then the particle methods that drop the assumption when the density stops being an ellipse. A robot odometer's uncertainty is a two-by-two covariance drifting exactly as the draggable one here.

Further reading

The references below treat the multivariate Gaussian as a closure result first and a formula second. If you take away one thing, take away the picture of an ellipse whose slices all have the same width, and the recollection that every standard operation keeps you inside the family.

Cheat sheet

TermMeaning here
Joint density$f(x)=\frac{1}{(2\pi)^{n/2}\sqrt{\det\Sigma}}\exp\!\big(-\tfrac12(x-\mu)^{\mathsf T}\Sigma^{-1}(x-\mu)\big)$
Quadratic form$(x-\mu)^{\mathsf T}\Sigma^{-1}(x-\mu)$, the squared Mahalanobis distance; level sets are ellipsoids
Covariance ellipseImage of the unit circle under $L$ where $\Sigma=LL^{\mathsf T}$; its axes are the eigenvectors of $\Sigma$
MarginalStill Gaussian; the marginal over a subset uses the matching sub-block of $\Sigma$, so $\mathrm{Var}(x_1)=\Sigma_{11}$
Conditional mean$\mu_{2\mid 1}=\mu_2+\Sigma_{21}\Sigma_{11}^{-1}(a-\mu_1)$, linear in the conditioning value
Schur complement$\Sigma_{2\mid 1}=\Sigma_{22}-\Sigma_{21}\Sigma_{11}^{-1}\Sigma_{12}=\sigma_2^2(1-\rho^2)$; independent of $a$
Affine map$x\mapsto Ax+b$ sends $\mathcal{N}(\mu,\Sigma)$ to $\mathcal{N}(A\mu+b,\;A\Sigma A^{\mathsf T})$
Sampling$x=\mu+Lz$ with $\Sigma=LL^{\mathsf T}$ and $z\sim\mathcal{N}(0,I)$ — what Stats.sample.mvnormal does
8

Check your understanding

0/4 answered