Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

The reverse process

Walking downhill in probability, one noisy step at a time

The forward process defines a chain of distributions $q(x_t)$ that starts at the data and ends at noise. The reverse process tries to walk that chain backwards. At each step the network estimates the noise $\varepsilon_\theta(x_t,t)$, and that estimate converts into an estimate of the clean point, $\hat{x}_0 = (x_t - \sqrt{1-\bar\alpha_t}\,\varepsilon_\theta)/\sqrt{\bar\alpha_t}$. A small step is then taken toward that estimate, scaled down so it does not overshoot the noise that is still present.

Every sampler in this volume is a variation on that one move, and they differ in one question: how much randomness is re-injected at each step? Re-inject the full amount dictated by the forward process and you get the ancestral sampler; re-inject none and you get a deterministic ordinary differential equation; re-inject something in between and you get a whole family in between.

The canvas below shows the target distribution as a blue density — three Gaussian modes — with the sampler's trajectory drawn over it in pink. The black dot is the starting noise. Press Reseed for a different start, switch methods, and drag the step slider to watch the path get coarser or finer.

Blue shading is the target density. The big pink line is the sampled trajectory from the black start dot; the outlined dot is where it ends. Under DDPM, five fainter runs from the same start are drawn to show the spread.

The analytic oracle is the exact posterior mean of the toy mixture, so what you see is the sampler, not a network's error. Part 3 trained the network; this part isolates the loop.

💡 By the end of this part you'll be able to write the reverse update, say exactly what ancestral DDPM adds that DDIM removes, and reason about the trade between step count, sample quality and compute.
2

Ancestral DDPM

Add the noise back, on purpose

The ancestral sampler is the reverse process written down literally. At each step it computes the mean of $q(x_{t-1}\mid x_t)$ using the network's noise estimate, and then samples around that mean with the variance the forward process prescribes. In symbols, with $\beta_t$ the step's noise and $\bar\alpha$ the cumulative signal,

$$ x_{t-1} = \frac{1}{\sqrt{\alpha_t}}\Big(x_t - \frac{\beta_t}{\sqrt{1-\bar\alpha_t}}\,\varepsilon_\theta(x_t,t)\Big) + \sigma_t z, \qquad z\sim\mathcal{N}(0,I) $$

That final $\sigma_t z$ is the part people find strange. The model has just estimated the noise, so why put noise back? Because the goal is not to recover one particular clean point — that information is gone — but to sample from the conditional distribution. The true reverse distribution is not a spike at its mean; it has a spread, and ignoring that spread produces samples that are too smooth and too similar to one another. Injecting the prescribed variance is what keeps the sampler honest about its own uncertainty.

Switch the demo to DDPM and press Reseed a few times. The same starting noise gives a different trajectory and a different endpoint every time, and the faint runs on the canvas show why: the randomness at each step is a real degree of freedom. This is also the property that makes DDPM expensive in practice, because the noise must be integrated carefully across the whole trajectory for the sample to be valid.

⚠ Ancestral does not mean accurate. Injecting the correct variance makes the sampler a valid draw from the model distribution only if the network is perfect. With a learned model the injected noise can amplify errors, which is one reason deterministic samplers are usually preferred once quality matters more than diversity.
3

Deterministic DDIM

Set the variance to zero and lose nothing but the dice

DDIM is what happens when you notice that the reverse process does not have to be Markovian or stochastic. The same network, evaluated at a sparse set of timesteps, can define a deterministic map from noise to sample:

$$ x_{t-1} = \sqrt{\bar\alpha_{t-1}}\,\hat{x}_0 + \sqrt{1-\bar\alpha_{t-1}}\,\varepsilon_\theta(x_t,t) $$

Read it as two endpoints: reconstruct the clean point $\hat{x}_0$, then re-noise it to the next noise level using the network's own noise estimate. There is no $\sigma_t z$, so the whole trajectory is fixed once the starting noise is fixed. Run the sampler twice with the same seed and you get bit-identical output, which is why DDIM is the usual choice for interpolation, inversion and editing: the map is a function, and functions can be differentiated and inverted.

The histogram below is the experiment that makes the difference unmistakable. It takes one fixed starting noise and runs the reverse process 120 times, recording the final $x$-coordinate each time. Under DDIM every run follows the same path, so all 120 land on the same value and the histogram is a single spike. Under DDPM each run draws fresh noise at every step, so the endpoints scatter around the mode — the same fixed start has become a distribution.

Final $x$-coordinate over 120 runs from one fixed starting point, using the method and step count selected above.

DDIM concentrates all 120 runs into one bar because it is deterministic; DDPM spreads them into a bell because it is not. Neither is wrong — they are two answers to "how much of the reverse process should be randomness?"

The spread canvas recomputes when you change the method or the step count, so you can also see DDPM's spread shrink as the steps get finer.

A useful way to hold the two samplers in your head: DDPM integrates a stochastic differential equation with the noise left in, and DDIM integrates the probability-flow ODE that has the same marginals. They agree in the limit of infinitely many steps and diverge in how they spend a finite budget, and both are special cases of the $\eta$-parameterised family where $\eta$ interpolates between them.

4

Step count and what it costs

The slider everyone wants to move left

Training touches a thousand noise levels; sampling can use far fewer because DDIM does not need to follow the chain one step at a time. Drag the slider down to ten steps and the trajectory becomes a jagged polyline that hops between noise levels; drag it to two hundred and it hugs a smooth curve. The sample quality follows the same pattern: too few steps and the path cuts corners, missing modes or overshooting them, while past a point extra steps cost compute and change almost nothing.

This is the central engineering trade of diffusion inference. Every step is one forward pass of the network, so halving the steps halves the latency, and the entire distillation literature exists to buy back the quality lost. The number that matters is not the count of steps but the size of the jump each step takes in noise level, which is why schedules with rapidly changing SNR need denser sampling there.

There is a subtlety the demo makes visible. Because DDIM is deterministic, a coarse run does not average out its own error over many stochastic steps — it commits to a path, and if that path is wrong it stays wrong. Ancestral DDPM has more chances to correct itself but pays in noise. Practical systems often use a deterministic sampler with a good schedule and accept a modest step count, because the wall-clock is what users feel.

⚠ Steps are not free even when the canvas is small. At 512×512 pixels with a U-Net, one step is a few hundred milliseconds on a datacentre GPU, so the difference between 50 and 20 steps is a visible wait. The toy here is instant; the reasoning about the trade is the same.
5

Where this shows up

Sampler choice is a product decision

Latent diffusion

Sampling in a smaller space

Most image systems run this loop in a compressed latent rather than in pixels, so one step costs far less and the step-count trade moves in the user's favour. The later parts of this volume build that compressor.

Serving

Latency is the whole game

A generation request is a long-running job whose cost is steps times per-step latency. The LLM serving guide reasons about batching, queues and tail latency for exactly this kind of workload; the sampling loop is the compute kernel it schedules.

Two threads lead out of this part. The first is the score function, which explains why the network's noise estimate is also a gradient and why the deterministic and stochastic samplers describe the same object; the second is the machinery of conditioning, which turns an unconditional sampler into one that follows a prompt. The next part takes the first thread.

Further reading

The sampling literature is where diffusion stopped being slow. These four papers are the arc: the ancestral sampler, the deterministic one, the unification, and the distillation that made interactive generation possible.

If you read one, read DDIM, because it is the paper that reframed sampling as an ODE and every fast sampler since has been a variation on that idea.

Cheat sheet

TermMeaning here
Reverse processIteratively denoise from $x_T$ down to $x_0$
$\hat{x}_0$Current estimate of the clean point, from $x_t$ and $\varepsilon_\theta$
DDPMAncestral: add variance $\sigma_t z$ at every step; stochastic
DDIMDeterministic: no added noise; same seed gives the same sample
$\eta$ familyInterpolates between DDIM ($\eta=0$) and DDPM ($\eta=1$)
Step countForward passes at generation time; the main latency dial
Probability-flow ODEThe deterministic equation whose marginals match the forward chain
6

Check your understanding

0/4 answered