Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

The denoising objective

Ask the network to name the noise

Take a clean point $x_0$, pick a random timestep $t$, and use the closed form from the previous part to jump straight to $x_t = \sqrt{\bar\alpha_t}\,x_0 + \sqrt{1-\bar\alpha_t}\,\varepsilon$. You know $\varepsilon$ exactly, because you generated it. So the training objective can be as blunt as this: given $x_t$ and $t$, predict $\varepsilon$, and penalise the squared error.

$$ \mathcal{L}(\theta) = \mathbb{E}_{x_0,\,t,\,\varepsilon}\!\left[\, \lVert \varepsilon - \varepsilon_\theta(x_t, t) \rVert^2 \,\right] $$

This is the simplified DDPM loss, and its elegance is that a hard probabilistic problem has become ordinary regression. There is no adversarial game and no likelihood to estimate directly: draw a batch, jump to a random noise level, and take a gradient step. The reason the shortcut is legitimate is that predicting the noise that was added is equivalent to estimating the score of the noisy distribution up to a known factor, a fact the fifth part makes precise.

The network below does not know about images. It is a tiny multilayer perceptron with two inputs, two tanh hidden layers of sixteen units, and two outputs, trained with hand-written backpropagation on a toy distribution — three Gaussian blobs in the plane. Everything true of this network is true of a billion-parameter U-Net; the only difference is scale.

💡 By the end of this part you'll be able to state the denoising objective, explain what the network is asked to output at each noise level, and say why ε-, x₀- and v-prediction are three parameterisations of the same information rather than three different models.
2

Train a denoiser in the browser

A real network, a real loss curve, one slider

The grey dots are the toy dataset: three clusters, roughly two hundred points. The blue dots are those points after the forward process at the noise level the slider selects, so at low $t$ they sit on the data and at high $t$ they are scattered almost at random. The pink arrows are the network's denoising attempt: each arrow starts at a noisy point and ends at the clean point the network believes it came from.

Press Retrain and watch the loss curve grow from right to left as the epochs tick past. At the start the network is guessing and the error is close to what you get from predicting zero. By the end the arrows point convincingly at the clusters. Drag the slider while the network is fixed and you will see it get worse — it was trained for one noise level, and a different one is out of distribution until you retrain.

Top: the toy data in grey, the noisy points at the selected level in blue, and the network's denoised estimate as pink arrows. Bottom: the training loss, one point per epoch.

The network is seeded, so the same slider value and the same retrain count always produce the same curve. Nothing here uses wall-clock randomness.

A full run is 300 epochs over 160 points — small enough to finish while you read this sentence, large enough to visibly learn.

⚠ The loss floor depends on $t$. At high signal-to-noise ratio $\sqrt{1-\bar\alpha_t}$ is tiny, so the noise is almost invisible in $x_t$ and even a perfect model cannot do much better than guessing zero. At low signal-to-noise ratio the network can nearly read $\varepsilon$ off $x_t$. A flat-looking curve at small $t$ is not a broken network; it is an under-determined problem.
3

Epsilon, x₀ and v

Three ways to say the same thing

Given $x_t$ and $t$, the quantities $\varepsilon$, $x_0$ and $v$ are algebraically interchangeable. The forward equation is one linear relation between them, so once you know any one of the three you can recover the others exactly:

The demo's network predicts $\varepsilon$, which is why the arrows are computed by first converting its output to $x_0$. You can see that conversion in the readout: the same network output, expressed three ways.

TargetFormulaTarget variance across $t$
Noise $\varepsilon$$(x_t - \sqrt{\bar\alpha_t}x_0)/\sqrt{1-\bar\alpha_t}$Constant: unit normal at every step
Clean $x_0$$(x_t - \sqrt{1-\bar\alpha_t}\,\varepsilon)/\sqrt{\bar\alpha_t}$Data variance, blown up as $\bar\alpha_t\to0$
Velocity $v$$\sqrt{\bar\alpha_t}\,\varepsilon - \sqrt{1-\bar\alpha_t}\,x_0$Roughly constant across the whole schedule
4

Why the parameterisation matters

Same information, very different gradients

If the three targets were exactly recoverable from one another, the choice would be cosmetic. It is not, because a network is a finite approximation trained with finite precision, and the loss weights the different error scales very differently. Training on $x_0$ near the end of the schedule means multiplying the target by $1/\sqrt{\bar\alpha_t}$, which can be a hundredfold amplification of the label: a tiny absolute error becomes a huge loss, gradients explode, and the model spends its capacity on timesteps where the clean image is unrecoverable anyway.

$\varepsilon$-prediction sidesteps that by keeping the target at unit variance everywhere, at the cost of the opposite pathology at high SNR, where the noise is nearly invisible and the model's output is dominated by the prior mean. $v$-prediction splits the difference and was introduced precisely to keep the target's scale stable at both ends of the schedule; it is the default in several of the systems that produce images today.

All three can be made to work in the same architecture, and they often are: a single network can emit all three heads, and a sampler can pick whichever is numerically convenient. What changes is not what the network knows, but how hard it is to learn it. This is a general lesson about diffusion that repeats everywhere — the mathematics is elegant and the numerics decide what actually trains.

⚠ The network above is trained at one noise level. A real denoiser is conditioned on $t$ as well, usually through a sinusoidal embedding injected into every residual block, so one set of weights serves the whole schedule. Everything you see here happens inside a real denoiser too; it just happens a thousand times, at a thousand noise levels, with the same weights.
5

Where this shows up

Denoising is gradient estimation

Optimisation

Gradient descent with a schedule

The training loop here is plain stochastic gradient descent with a hand-written backward pass. The nonlinear optimisation guide covers why the learning rate and the loss scale interact the way they do — the same reason $\varepsilon$-prediction is easier to train than $x_0$-prediction.

Probability

Regression to a conditional mean

Minimising squared error makes the network converge to $\mathbb{E}[\varepsilon\mid x_t]$, the posterior mean. The linear regression chapter develops the same fact for ordinary least squares: the cheapest thing to learn is always an expectation.

Once a denoiser exists, sampling is a loop: start from noise, ask the network for its noise estimate, subtract a fraction of it, and repeat. The next part runs that loop two ways — one ancestral and stochastic, one deterministic — and counts what each costs.

Further reading

The denoising objective, and the parameterisation debate around it, are the subject of a small literature that is worth reading in order. The DDPM paper defines the loss; the v-prediction paper diagnoses its failure modes; and the progressive-distillation work is the first place the sampler's step count becomes a design variable.

If you read one thing, read the v-prediction paper's analysis of signal-to-noise truncation, because it explains why the same model can look excellent on one benchmark and broken on another.

Cheat sheet

TermMeaning here
Denoising loss$\mathbb{E}\lVert\varepsilon-\varepsilon_\theta(x_t,t)\rVert^2$; ordinary regression
$\varepsilon$-predictionPredict the noise; unit-variance target at every $t$
$x_0$-predictionPredict the clean point; target scale blows up near $t=T$
$v$-predictionPredict $\sqrt{\bar\alpha_t}\varepsilon-\sqrt{1-\bar\alpha_t}x_0$; stable scale
Conditioning on $t$One network serves all noise levels via a timestep embedding
MLPThe toy network here: 2→16→16→2, tanh, linear output
Loss floorSet by how much of $\varepsilon$ is recoverable from $x_t$ at that $t$
6

Check your understanding

0/4 answered