Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

The two families

Ancestral noise or a deterministic integration

Every sampler here runs the same reverse loop. At each step it asks the network for an estimate of the noise $\varepsilon$ left in the current sample and then takes a step toward less noise. The families differ in what the step looks like. The ancestral sampler — DDPM — treats the reverse process as a chain of noisy transitions: at every step it computes a mean and then adds a fresh draw of Gaussian noise whose variance is fixed by the schedule. Two runs from the same start give two different images, and the stochasticity is not a defect; it is the Markov chain doing exactly what it was trained to do.

The deterministic family throws the injected noise away. DDIM replaces the noisy transition with a straight interpolation toward the predicted clean sample, so the whole trajectory becomes a function of the starting noise: one start, one endpoint. The same rule can be read as an ordinary differential equation, the probability-flow ODE, whose velocity field is built from the denoiser. Once generation is an ODE, the entire numerical-analysis toolkit applies. Euler is the crudest integrator, Heun corrects Euler with a second field evaluation, and the multistep methods reuse predictions from earlier steps to buy accuracy at a fixed number of network calls.

💡 By the end of this part you'll be able to say why a deterministic sampler is a function and an ancestral one is a distribution, why a higher-order solver leaves less error at the same step count, and why the spacing of the timesteps is as much a design choice as the solver itself.
2

Solver order and error

What a second evaluation buys

The plot below integrates the probability-flow ODE on a linear schedule from a known start toward a known clean point $x_0 = 1$, and measures how far the endpoint lands from that point. The predictor is deliberately imperfect: it is an oracle model with its noise estimate inflated by $15\%$ whenever the signal has mostly been destroyed, which is the garden-variety failure of a real network at high noise. That single flaw is enough to make the solver's job non-trivial, and it is what the curves measure.

Euler is first order: halve the step size and the error roughly halves. Heun is second order: it takes an Euler step, evaluates the field again at the provisional endpoint, and averages the two, which costs twice as many network calls per step but makes the error fall much faster as the steps multiply. The dashed line is the floor the error would reach with a perfect predictor, computed by Diffusion.solverError; no solver can go below it.

Endpoint error against the number of steps, log scale. Solid curves use the imperfect predictor; the dashed curve is the perfect-model floor.

⚠ Heun is not free. Two field evaluations per step double the number of forward passes through the network, and the network is almost the whole cost of sampling. A second-order solver that halves the error at 16 steps is really spending 32 evaluations to do it, which is often the same as 32 Euler steps. The right comparison is error per network call, not error per step.
3

Timestep spacing

Where the steps land on the sigma curve

The schedule fixes how much noise each timestep carries. It is convenient to measure that as a noise-to-signal ratio, $\sigma_t = \sqrt{(1-\bar\alpha_t)/\bar\alpha_t}$, which falls from a large value at the start of sampling to nearly zero at the end. Uniform spacing — the familiar linspace over timesteps — takes equal steps along the index, but the index is not where the difficulty lives. Almost all of the structural decision-making happens when $\sigma$ is small and the image is nearly formed, and uniform spacing spends most of its budget at high $\sigma$ where there is little to decide.

Karras and coauthors warp the grid in $\sigma$-space instead, pulling steps densely toward small $\sigma$ and sparsely toward the noise end. The plot makes the difference visible: the marker sets are the same number of steps, but the Karras set leaves the smooth $\sigma$ curve and clusters at the low-noise end. With a fixed step budget, that reallocation is usually worth more than a higher-order solver, which is why every production sampler exposes a spacing knob beside its solver knob.

The noise-to-signal curve (solid) with the steps chosen by uniform spacing and by Karras spacing. Both use the step count from the solver demo above.

Larger ρ concentrates steps even harder at low σ. Uniform spacing is the special case where every step covers the same index interval.

💡 The curve is not an accident of the schedule: it is determined by $\bar\alpha_t$, the same quantity the forward process accumulated. Changing the schedule changes the shape, and changing the spacing changes which part of the shape the solver is allowed to resolve.
4

The step budget

The twenty-to-fifty sweet spot

The training process used a thousand forward steps, but sampling never has to. The reverse loop is far more forgiving than the forward chain, and the honest way to state the trade is in network evaluations: wall-clock time is roughly the number of forward passes times the cost of one pass. The bars below hold the solver fixed and spend four different budgets, from four steps to fifty, on the same trajectory.

In practice the curve flattens somewhere in the twenties. Below that the image is visibly under-resolved and the sampler has not had enough looks at the low-noise end; above fifty the extra evaluations buy a difference no one can see. That flat region is why the default sampler in a shipping pipeline is typically twenty to fifty steps with a second-order or multistep solver and a warped grid, and why the few-step methods in the next part are interesting at all: they try to move the knee of this curve down to one or four.

Endpoint error at 4, 8, 20 and 50 steps for the solver selected above. The highlighted bar is the current budget.

Multistep solvers such as DPM-Solver++ reach the flat region in roughly ten to twenty calls by reusing earlier predictions as extra terms in the update, rather than by evaluating the field again.

5

Where this shows up

The knobs every pipeline exposes

Image

Stable Diffusion samplers

The sampler menu in every Stable Diffusion interface is this part's taxonomy: ancestral samplers that inject noise, DDIM and its relatives that do not, and DPM-Solver++ variants that add history terms. The reverse-loop chapter builds the same loop on a real trained denoiser.

Flow & real time

Few-step and interactive generation

A straight-path flow-matching field can be integrated in a handful of Euler steps, which is what makes interactive editing and robot policies at control rate feasible. The ODE guide develops Euler and Runge–Kutta error orders, which is the analysis behind every curve on this page.

The same arithmetic organises video, audio and 3-D generation, where each network call is even more expensive and the step budget is measured in seconds of latency rather than milliseconds. Whenever a paper reports a sampling time, read it as solver, spacing and step count together — the three are chosen jointly and usually tuned against each other.

Further reading

Sampling is where diffusion stops being a probabilistic model and starts being numerical analysis. These are the papers that made that shift explicit: the one that removed the injected noise, the one that added higher-order integration, and the one that treated the schedule as a warped grid in noise-to-signal space.

Song and coauthors show that the reverse process has a deterministic counterpart; Lu and coauthors bring the multistep ODE solvers of numerical analysis into the loop; Karras and coauthors rework the whole sampler around the sigma parameterisation; and Ho and coauthors are the ancestral baseline everything else is measured against.

Cheat sheet

TermMeaning here
Ancestral samplerInjects fresh Gaussian noise at every step; DDPM; stochastic output
Deterministic samplerNo injected noise; the endpoint is a function of the starting noise
Probability-flow ODEThe differential equation whose solution is the deterministic reverse path
Euler stepFirst order: follow the field at the current point; $x \leftarrow x + v\,\Delta t$
Heun stepSecond order: average the field at the start and at a provisional endpoint; two evaluations
NFENumber of function evaluations; the true cost of a sampler
$\sigma_t$Noise-to-signal ratio $\sqrt{(1-\bar\alpha_t)/\bar\alpha_t}$; what the schedule really controls
Karras spacingNon-uniform grid, dense at low $\sigma$; Diffusion.karrasSigmas
Sweet spotRoughly 20–50 evaluations for images before the error curve flattens
7

Check your understanding

0/4 answered