Calculus of probability
A probability density is not a probability — it is a function you integrate to get one. That single fact makes the whole of this part calculus: probabilities are integrals, means and variances are integrals, and when a random variable is pushed through a map its density is divided by the derivative of that map. Sample a Gaussian with a fixed seed, push it through affine and nonlinear maps, and watch the density transform exactly as the change-of-variables formula predicts.
A density is a function you integrate
Probability as area under a curve
A random variable $X$ has a probability density $p(x)$ if probabilities are given by integrals of it:
The density itself is not a probability: $p(x)$ can exceed $1$ (a narrow Gaussian does), because it is a probability per unit length. Only its integral is a probability. The total area is fixed at $1$ — that is the normalisation condition that makes the whole thing a distribution. The same integral defines every expectation you will meet later:
Below is a standard Gaussian. Drag the two handles (or the sliders) and watch the shaded area, computed by a fine trapezoid rule, report the probability of the interval. The readout also integrates the whole line to confirm the total is $1$.
A standard Gaussian density. The shaded region is the trapezoid approximation of $\int_a^b p(x)\,dx$; drag the round handles to move the limits.
Mean, variance and the affine map
Two integrals, then a shift and a stretch
The two numbers that summarise a distribution are themselves integrals:
Because the integral is linear, an affine map moves the mean and scales the spread in the obvious way. If $Z \sim \mathcal{N}(0,1)$ then $Y = \mu + \sigma Z$ is Gaussian with
The demo samples $Z$ once with a fixed seed, forms $Y = \mu + \sigma Z$, and histograms both. The analytic densities are overlaid; the sample mean and sample standard deviation track $\mu$ and $\sigma$. Note that the sample standard deviation is the $\sigma$ scaling made visible.
Top: the seeded standard-normal draws. Bottom: the same draws after $y = \mu + \sigma z$, with the analytic $\mathcal{N}(\mu,\sigma^2)$ density overlaid.
Change of variables for densities
The density is divided by the derivative
An affine map is the easy case because its derivative is constant. For a general strictly monotone map $Y = g(Z)$, a small interval $dz$ is stretched into an interval of length $|g'(z)|\,dz$, so the probability $p_Z(z)\,dz$ must be packed into $p_Y(y)\,dy$. Setting the two probabilities equal gives the change-of-variables formula for densities:
The factor $|g'|$ is the one-dimensional Jacobian — exactly the local length scale from Part 5, where the two-dimensional version is $|\det J|$. Where the map stretches space, the density is thinned out; where it compresses space, the density piles up.
Pick a map, slide the probe, and read the three numbers that make the formula concrete: the density of $z$, the Jacobian factor $|g'(z)|$ at the probe, and their ratio, the density of $y$. The histogram is the empirical density of the transformed samples; the coloured curve is the analytic prediction.
Histogram of $y = g(z)$ for seeded Gaussian $z$, against the analytic $p_Y$ from the change-of-variables formula. The probe line marks a chosen $z$.
Reparameterisation
Sample once, move the parameters
There is a way to read the affine rule that turns out to be the workhorse of modern generative modelling. Fix the noise: draw $z \sim \mathcal{N}(0,1)$ once, and write every Gaussian draw as a deterministic function of that noise,
Now $\mu$ and $\sigma$ are ordinary parameters sitting outside the random draw, so the sample $y$ depends on them smoothly and differentiably. This is the reparameterisation trick: a gradient can flow through $\mu$ and $\sigma$ into the samples. Drag the sliders and the whole cloud of fixed $z$'s slides and stretches rigidly — no resampling, no jitter, just an affine map of a fixed set of points.
Top row: fixed seeded standard-normal draws $z$. Bottom row: the same draws after $y = \mu + \sigma z$. The cross marks each sample mean.
KL divergence
An integral of a density ratio
To measure how far one distribution $q$ is from another $p$, integrate the log density ratio against $q$:
The two forms are the same statement: the integral is an expectation under $q$, so it can be estimated by sampling from $q$. For two Gaussians it has a closed form,
so we can check the Monte-Carlo estimate against an exact answer. The demo shades the integrand $q(x)\log\frac{q(x)}{p(x)}$ — green where it is positive, magenta where negative — and reports both numbers. The area under that curve is the KL.
This is the quantity inside the training objective. Cross-entropy splits as $H(q,p) = H(q) + D_{\mathrm{KL}}(q\|p)$, and the negative log-likelihood a model minimises is cross-entropy with $q$ the data distribution; when $q$ is a posterior over latents, the KL term is the regulariser that keeps it near a prior.
Top: the two Gaussian densities $q$ and $p$. Bottom: the integrand $q\log(q/p)$, shaded with sign; its signed area is $D_{\mathrm{KL}}(q\|p)$.
Where this shows up
Every loss is an integral in disguise
Training a language model is this part applied at scale. The cross-entropy loss is $\mathbb{E}_{x\sim\text{data}}[-\log p_\theta(x)]$ — an expectation, hence an integral, hence the thing you differentiate; LLM Training builds its whole objective out of exactly that, and the KL regulariser in preference tuning is the divergence defined above. When the density lives in several variables the integral is a multiple integral and the Jacobian is a determinant — that is Part 6: Multiple integrals & coordinates. And when the density is constrained to a curved space, the Gaussian is written in the tangent plane and the change-of-variables factor becomes a metric — the beginning of Part 13: Calculus on manifolds.
Notation to carry forward
| Notation | Reads as | Where it comes up |
|---|---|---|
p(x) | The probability density — a non-negative function integrating to 1 | This part; every probabilistic model |
∫ₐᵇ p(x) dx | The probability that X lands in the interval | This part, Step 1 |
E[f(X)] = ∫ f p dx | The expectation as an integral against the density | Means, variances, losses, entropy |
μ, σ² | Mean and variance, both defined as integrals | Every Gaussian; summary statistics |
p_Y(y) = p_Z(g⁻¹(y)) / |g'(y)| | The change-of-variables rule for a monotone map | Step 3; Part 5 in one dimension |
|g'(z)|, |det J| | The local Jacobian factor dividing a transformed density | Part 5; Part 6 in higher dimensions |
y = μ + σz | Reparameterisation: fixed noise, differentiable parameters | VAEs; Step 4 |
D_KL(q‖p) = ∫ q log(q/p) dx | Expected log density ratio — an integral and an expectation | Step 5; LLM training objectives |
Further reading
- MIT 18.05, Introduction to Probability and Statistics — densities, expectations and transformations done carefully.
- Seeing Theory, Brown University — visual introductions to distributions, expectation and the central limit theorem.
- Gregory Gundersen, The reparameterization trick — why sampling can be made differentiable, with the VAE objective in view.