Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

A density is a function you integrate

Probability as area under a curve

A random variable $X$ has a probability density $p(x)$ if probabilities are given by integrals of it:

$$ \Pr(a < X < b) \;=\; \int_a^b p(x)\,dx, \qquad p(x) \ge 0, \qquad \int_{-\infty}^{\infty} p(x)\,dx = 1. $$

The density itself is not a probability: $p(x)$ can exceed $1$ (a narrow Gaussian does), because it is a probability per unit length. Only its integral is a probability. The total area is fixed at $1$ — that is the normalisation condition that makes the whole thing a distribution. The same integral defines every expectation you will meet later:

$$ \mathbb{E}[f(X)] \;=\; \int_{-\infty}^{\infty} f(x)\,p(x)\,dx. $$

Below is a standard Gaussian. Drag the two handles (or the sliders) and watch the shaded area, computed by a fine trapezoid rule, report the probability of the interval. The readout also integrates the whole line to confirm the total is $1$.

A standard Gaussian density. The shaded region is the trapezoid approximation of $\int_a^b p(x)\,dx$; drag the round handles to move the limits.

💡 The density is the thing you integrate. Every quantity in probability — probability, mean, variance, entropy, KL — is an integral of some function against a density. That is why this volume's integration parts keep returning.
2

Mean, variance and the affine map

Two integrals, then a shift and a stretch

The two numbers that summarise a distribution are themselves integrals:

$$ \mu \;=\; \mathbb{E}[X] \;=\; \int x\,p(x)\,dx, \qquad \sigma^2 \;=\; \mathbb{E}\big[(X-\mu)^2\big] \;=\; \int (x-\mu)^2\,p(x)\,dx. $$

Because the integral is linear, an affine map moves the mean and scales the spread in the obvious way. If $Z \sim \mathcal{N}(0,1)$ then $Y = \mu + \sigma Z$ is Gaussian with

$$ \mathbb{E}[Y] = \mu + \sigma\,\mathbb{E}[Z] = \mu, \qquad \operatorname{Var}(Y) = \sigma^2\operatorname{Var}(Z) = \sigma^2. $$

The demo samples $Z$ once with a fixed seed, forms $Y = \mu + \sigma Z$, and histograms both. The analytic densities are overlaid; the sample mean and sample standard deviation track $\mu$ and $\sigma$. Note that the sample standard deviation is the $\sigma$ scaling made visible.

Top: the seeded standard-normal draws. Bottom: the same draws after $y = \mu + \sigma z$, with the analytic $\mathcal{N}(\mu,\sigma^2)$ density overlaid.

3

Change of variables for densities

The density is divided by the derivative

An affine map is the easy case because its derivative is constant. For a general strictly monotone map $Y = g(Z)$, a small interval $dz$ is stretched into an interval of length $|g'(z)|\,dz$, so the probability $p_Z(z)\,dz$ must be packed into $p_Y(y)\,dy$. Setting the two probabilities equal gives the change-of-variables formula for densities:

$$ p_Y(y) \;=\; \frac{p_Z\big(g^{-1}(y)\big)}{\big|g'(y)\big|}, \qquad\text{equivalently}\qquad p_Y\big(g(z)\big) \;=\; \frac{p_Z(z)}{\big|g'(z)\big|}. $$

The factor $|g'|$ is the one-dimensional Jacobian — exactly the local length scale from Part 5, where the two-dimensional version is $|\det J|$. Where the map stretches space, the density is thinned out; where it compresses space, the density piles up.

Pick a map, slide the probe, and read the three numbers that make the formula concrete: the density of $z$, the Jacobian factor $|g'(z)|$ at the probe, and their ratio, the density of $y$. The histogram is the empirical density of the transformed samples; the coloured curve is the analytic prediction.

Histogram of $y = g(z)$ for seeded Gaussian $z$, against the analytic $p_Y$ from the change-of-variables formula. The probe line marks a chosen $z$.

⚠️ At $z=0$ the map $z^3$ has $g'(z)=0$, so the formula predicts an infinite density there — the samples pile into a spike. A map whose derivative vanishes is not a good reparameterisation: it is not invertible locally.
4

Reparameterisation

Sample once, move the parameters

There is a way to read the affine rule that turns out to be the workhorse of modern generative modelling. Fix the noise: draw $z \sim \mathcal{N}(0,1)$ once, and write every Gaussian draw as a deterministic function of that noise,

$$ y \;=\; \mu + \sigma z, \qquad z \sim \mathcal{N}(0,1). $$

Now $\mu$ and $\sigma$ are ordinary parameters sitting outside the random draw, so the sample $y$ depends on them smoothly and differentiably. This is the reparameterisation trick: a gradient can flow through $\mu$ and $\sigma$ into the samples. Drag the sliders and the whole cloud of fixed $z$'s slides and stretches rigidly — no resampling, no jitter, just an affine map of a fixed set of points.

Top row: fixed seeded standard-normal draws $z$. Bottom row: the same draws after $y = \mu + \sigma z$. The cross marks each sample mean.

💡 Why this matters: a variational autoencoder maximises an objective containing $\mathbb{E}_{q}[\,\cdot\,]$ over a Gaussian $q$. Without reparameterisation the expectation is not differentiable in the encoder's parameters; with it, one seeded noise draw carries the gradient. This is Part 9's chain rule doing the work.
5

KL divergence

An integral of a density ratio

To measure how far one distribution $q$ is from another $p$, integrate the log density ratio against $q$:

$$ D_{\mathrm{KL}}(q \,\|\, p) \;=\; \int q(x)\,\log\frac{q(x)}{p(x)}\,dx \;=\; \mathbb{E}_{x \sim q}\!\left[\log\frac{q(x)}{p(x)}\right]. $$

The two forms are the same statement: the integral is an expectation under $q$, so it can be estimated by sampling from $q$. For two Gaussians it has a closed form,

$$ D_{\mathrm{KL}}\big(\mathcal{N}(\mu_1,\sigma_1^2)\,\|\,\mathcal{N}(\mu_2,\sigma_2^2)\big) \;=\; \log\frac{\sigma_2}{\sigma_1} \;+\; \frac{\sigma_1^2 + (\mu_1-\mu_2)^2}{2\sigma_2^2} \;-\; \frac{1}{2}, $$

so we can check the Monte-Carlo estimate against an exact answer. The demo shades the integrand $q(x)\log\frac{q(x)}{p(x)}$ — green where it is positive, magenta where negative — and reports both numbers. The area under that curve is the KL.

This is the quantity inside the training objective. Cross-entropy splits as $H(q,p) = H(q) + D_{\mathrm{KL}}(q\|p)$, and the negative log-likelihood a model minimises is cross-entropy with $q$ the data distribution; when $q$ is a posterior over latents, the KL term is the regulariser that keeps it near a prior.

Top: the two Gaussian densities $q$ and $p$. Bottom: the integrand $q\log(q/p)$, shaded with sign; its signed area is $D_{\mathrm{KL}}(q\|p)$.

q density p density integrand positive integrand negative
6

Where this shows up

Every loss is an integral in disguise

Training a language model is this part applied at scale. The cross-entropy loss is $\mathbb{E}_{x\sim\text{data}}[-\log p_\theta(x)]$ — an expectation, hence an integral, hence the thing you differentiate; LLM Training builds its whole objective out of exactly that, and the KL regulariser in preference tuning is the divergence defined above. When the density lives in several variables the integral is a multiple integral and the Jacobian is a determinant — that is Part 6: Multiple integrals & coordinates. And when the density is constrained to a curved space, the Gaussian is written in the tangent plane and the change-of-variables factor becomes a metric — the beginning of Part 13: Calculus on manifolds.

7

Notation to carry forward

NotationReads asWhere it comes up
p(x)The probability density — a non-negative function integrating to 1This part; every probabilistic model
∫ₐᵇ p(x) dxThe probability that X lands in the intervalThis part, Step 1
E[f(X)] = ∫ f p dxThe expectation as an integral against the densityMeans, variances, losses, entropy
μ, σ²Mean and variance, both defined as integralsEvery Gaussian; summary statistics
p_Y(y) = p_Z(g⁻¹(y)) / |g'(y)|The change-of-variables rule for a monotone mapStep 3; Part 5 in one dimension
|g'(z)|, |det J|The local Jacobian factor dividing a transformed densityPart 5; Part 6 in higher dimensions
y = μ + σzReparameterisation: fixed noise, differentiable parametersVAEs; Step 4
D_KL(q‖p) = ∫ q log(q/p) dxExpected log density ratio — an integral and an expectationStep 5; LLM training objectives
8

Further reading

9

Check your understanding

0/4 answered