Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

What the score is

The gradient of log density, and why anyone would learn it

For a density $p(x)$, the score is the vector field $\nabla_x \log p(x)$. It points in the direction in which the log density increases fastest, so following it climbs toward the modes. For a single Gaussian $\mathcal{N}(\mu, \sigma^2)$ there is nothing to compute — the score is $-(x-\mu)/\sigma^2$, a field of arrows all pointing straight at the mean with magnitude growing linearly with distance.

The reason score gets its own name is that it is often easier to estimate than the density itself. The density needs a normalising constant, which is an integral over everything and is generally intractable; the score does not, because the gradient of $\log Z$ is zero. That single observation is the foundation of score-based generative modelling: fit the gradient, never compute the constant, and never evaluate $p$ directly.

The difficulty is that in regions where the data are sparse, the score is badly estimated and the direction it gives is unreliable. The fix is to add noise: a noisy density has a smoother, better-behaved score, and if you anneal the noise down from large to small, the collection of scores at every level lets you walk from pure noise to data. That is exactly what the forward process was doing, and it is exactly what the network in part 3 was learning.

💡 By the end of this part you'll be able to read a score field, run Langevin dynamics by dragging a point, state the exact relation $\nabla\log p_t = -\varepsilon_\theta/\sqrt{1-\bar\alpha_t}$ between the denoiser and the score, and compare the VE and VP SDEs that formalise the whole picture.
2

Langevin dynamics

Drag the dot and watch it fall into a mode

If the score points uphill in probability, a point that repeatedly steps along it should find the modes. That is the idea of Langevin dynamics: take a gradient step, then add a little noise, and repeat. Written down, with a step size $\delta$,

$$ x_{k+1} = x_k + \delta\,\nabla_x\log p(x_k) + \sqrt{2\delta}\,z_k, \qquad z_k\sim\mathcal{N}(0,I) $$

The deterministic part climbs; the stochastic part stops the point from getting stuck at the top of a mode and lets it explore the whole distribution. The same $\sqrt{2\delta}$ balance that appears here reappears in the SDEs below, and it is not a liberty: with that coefficient, and only that coefficient, the stationary distribution of the walk is exactly $p$.

The demo draws the target density as blue shading and the score as a grid of arrows, coloured from grey (small) to blue (large). The pink dot is a starting point you can drag anywhere; the pink curve is 220 Langevin steps from that position. Drag it into a valley between two modes and watch it get pulled to one side by the field before the noise jitters it along.

Blue shading: the target density (three Gaussian modes). Arrows: the score field $\nabla\log p$. Pink curve: a Langevin walk from the draggable start dot.

Drag the pink dot anywhere in the frame. The walk is recomputed from that point with a fixed seed, so the same start always gives the same path.

⚠ Step size is a real trade-off. Too small and the walk crawls and never mixes; too large and each gradient step overshoots the mode and the walk flies off. The condition is roughly $\delta$ below the inverse of the largest curvature of $\log p$, and in high dimensions that is a serious constraint — which is why the noise-annealed multi-scale schedules exist in the first place.
3

The score is the denoiser

One identity ties this part to the last three

Here is the bridge. The forward process gives a noisy density $p_t$ at every level, and for the Gaussian perturbation $q(x_t\mid x_0) = \mathcal{N}(\sqrt{\bar\alpha_t}\,x_0, (1-\bar\alpha_t)I)$ the score of the noisy density has a closed relation to the noise that was added:

$$ \nabla_{x_t}\log p_t(x_t) = -\,\frac{\mathbb{E}[\varepsilon \mid x_t]}{\sqrt{1-\bar\alpha_t}} = -\,\frac{\varepsilon_\theta(x_t,t)}{\sqrt{1-\bar\alpha_t}} $$

So a network trained to predict noise is a network trained to predict the score, up to one scalar that depends only on $t$. Nothing in part 3 needs to change; it was doing score matching the whole time under a different name. This equivalence is the SMLD-to-DDPM bridge: the older score-matching line of work learned $\nabla\log p$ with denoising autoencoders at a ladder of noise levels, and the diffusion line learned $\varepsilon$ at a ladder of timesteps, and the two are the same object rescaled.

The small plot below makes the rescaled factor visible. The curve is $-1/\sqrt{1-\bar\alpha_t}$, the multiplier that converts a unit noise prediction into a score, plotted against $\bar\alpha_t$. At high signal the factor is near one, so the two are almost the same number; near the end of the schedule it diverges, and that divergence is exactly the reason a sampler trained on noise predictions struggles to use them as scores at very low signal-to-noise ratio.

The conversion factor $-1/\sqrt{1-\bar\alpha_t}$ as a function of $\bar\alpha_t$. The dashed line is the value selected by the slider.

A unit noise prediction $\varepsilon=1$ becomes a score of $-1/\sqrt{1-\bar\alpha_t}$: a small number at high signal, a large one at low signal.

This identity also explains the sampler from the previous part. DDIM's deterministic update is a discretised ordinary differential equation whose drift is built from this score; DDPM's stochastic update adds the Langevin-style noise that keeps the walk at the right temperature. The two samplers are two ways of integrating the same field.

4

VE versus VP SDEs

Two continuous-time limits of the same forward process

Take the number of forward steps to infinity and the discrete chain becomes a stochastic differential equation. Two parameterisations dominate the literature, and they differ in what they keep fixed as time runs.

Both define a valid bridge from data to noise; both admit a reverse-time SDE driven by the score; and a sampler can be built from either. The demo plots the two quantities side by side, normalised so their shapes are comparable: either the marginal standard deviation as a function of time, or the diffusion coefficient $g(t)$ itself. Switch modes and move the slider to see how the choice of maximum noise scale reshapes the VE curve.

VE in accent-2, VP in accent. Each curve is normalised to its own maximum so the shapes can be compared on one axis.

The two differ in more than shape. VE's score has magnitude proportional to $1/\sigma(t)$, so it explodes at small noise and needs careful weighting; VP's is bounded. That is why VP is the default for images and VE remains common for audio, where the dynamic range is enormous anyway.

A third parameterisation, the variance-preserving with a cosine schedule, is what most image models actually use; and flow matching, which the later parts of this volume cover, replaces the SDE entirely with a deterministic interpolation between noise and data. All of them are members of one family, and this part is the reason they can be compared at all.

5

Where this shows up

One field, many samplers

Probability in action

Langevin and MCMC

The walk you dragged is the same one that appears in Bayesian computation: Langevin and its Metropolis-adjusted cousin are workhorses of the Bayesian inference chapter. There the score comes from a likelihood; here it comes from a network, but the dynamics are identical.

Calculus

Gradient fields

A score is a gradient field, and the calculus in motion guide develops the machinery: gradient, divergence, and the continuous-time flows that the SDEs here describe. The reverse-time SDE is a Fokker–Planck problem in disguise.

This is the conceptual end of the first arc of the guide. The remaining parts take the same objects into practice: conditioning the score on a prompt, compressing the space so sampling is affordable, replacing the U-Net with a transformer, and extending the whole apparatus to video and audio. Every one of them assumes the picture on this page — a vector field you learn, and a walk you take through it.

Further reading

The score-based line is older than diffusion and repays reading on its own terms. These references cover the field, the sampler, and the unification with diffusion, in that order.

If you read one, read the SDE paper: it is the paper in which all of the pieces in this volume finally became one piece.

Cheat sheet

TermMeaning here
Score$\nabla_x\log p(x)$; points uphill in log density, needs no normalising constant
Gaussian score$-(x-\mu)/\sigma^2$
Langevin step$x+\delta\,\nabla\log p+\sqrt{2\delta}\,z$; climbs with noise added
Denoiser identity$\nabla\log p_t = -\varepsilon_\theta/\sqrt{1-\bar\alpha_t}$
VE SDEZero drift, exploding variance; $g=\sqrt{2\sigma\sigma'}$
VP SDEDrift $-\tfrac12\beta x$, diffusion $\sqrt{\beta}$; variance stays one
Reverse SDERun the forward SDE backwards with the score as the correction term
Noise annealingUse scores at many noise levels so the walk starts where the field is smooth
6

Check your understanding

0/4 answered