Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

What the forward process does

A small amount of noise, repeated many times

Start with a clean image $x_0$. Add a little Gaussian noise and you get $x_1$. Add a little more and you get $x_2$, and so on for a thousand steps until $x_T$ is indistinguishable from pure static. This chain is the forward process, and it is not learned — it is a fixed recipe. Each step is $x_t = \sqrt{1-\beta_t}\,x_{t-1} + \sqrt{\beta_t}\,\varepsilon_t$ with $\varepsilon_t \sim \mathcal{N}(0, I)$, where the coefficients are chosen so that the signal shrinks and the noise grows by matching amounts.

Two design choices are worth noticing now, because they are the reason everything later simplifies. The noise is standard Gaussian, so the whole chain stays in the family of Gaussian distributions; and the scaling by $\sqrt{1-\beta_t}$ keeps the variance of $x_t$ at roughly one, so the numbers never blow up or vanish. That second choice is the difference between diffusion and a random walk that merely gets larger.

Slide the timestep below and watch the blobs dissolve. The patch grid is drawn on top because a diffusion network does not see the image as pixels — it sees it as a grid of patches, and the noise is applied to all of them at once.

A procedurally generated 64×64 image at timestep $t$, with an 8×8 patch grid overlaid. The same noise field is reused at every $t$, so dragging the slider shows one image being destroyed, not a new image each time.

Switch schedules at the same $t$ and compare: the cosine curve is still recognisable where the linear one is already gone.

💡 By the end of this part you'll be able to write $q(x_t \mid x_0)$ in closed form, explain why one line replaces a thousand-step loop during training, and choose a schedule by reading its $\bar\alpha_t$ and SNR curves rather than by trial and error.
2

The closed form

Jump to any timestep in one line

The step-by-step chain is a definition, but training does not want to run it. If you fold the recursion, the sum of many small Gaussians is again a Gaussian, and you get the single most useful equation in the whole subject:

$$ q(x_t \mid x_0) = \mathcal{N}\!\left(\sqrt{\bar\alpha_t}\,x_0,\; (1-\bar\alpha_t) I\right) $$

where $\bar\alpha_t = \prod_{s=1}^{t}(1-\beta_s)$ is the cumulative product of the signal coefficients. In words: to produce $x_t$ you do not need steps $1$ through $t-1$; you just scale the clean image by $\sqrt{\bar\alpha_t}$, add Gaussian noise of standard deviation $\sqrt{1-\bar\alpha_t}$, and you are done. The mean is a shrunken copy of $x_0$ and the variance is the noise that has accumulated, and the two always sum to one.

The lower canvas shows why this matters. The grey histogram is the distribution of pixel values in the clean image; the pink one is the same pixels after the forward process at the current $t$. At small $t$ the pink bars sit on top of the grey ones, just a little wider. As $t$ grows the pink distribution flattens toward a standard normal, which is exactly the statement that $q(x_t \mid x_0)$ has forgotten $x_0$.

Pixel-value distributions: original ($x_0$) against noisy ($x_t$) at the timestep chosen above.

The forward process is a bridge between two distributions that both have a simple name. At $t=0$ it is a point mass at your data. At large $t$ it is standard normal noise. Everything a diffusion model does happens in between, and the closed form lets it sample any point on that bridge for free.

This is also why training is cheap: at every gradient step you draw a random $t$, jump straight to $x_t$, and regress. No rollout, no backpropagation through time.

⚠ The variance does not stay at one. The closed form keeps the variance of $x_t$ near one only if $x_0$ itself has unit variance. Real pixel data are in $[0,1]$, so the practical recipe rescales them (usually to $[-1,1]$) before the process starts. The algebra is unaffected; only the bookkeeping of constants is.
3

ᾱ, SNR and the schedules

Two curves that describe an entire chain

Everything you need to know about a noise schedule is in one function of time, $\bar\alpha_t$, and everything about how hard the denoising problem is at time $t$ is in the signal-to-noise ratio $\mathrm{SNR}(t) = \bar\alpha_t / (1-\bar\alpha_t)$. High SNR means the image is mostly intact and the model has an easy job; low SNR means the image is mostly noise and the model is guessing at a coarse scale. The crossover where SNR equals one is where the picture stops being visible by eye, and it lands at very different times for different schedules.

The first plot traces $\bar\alpha_t$ for the linear schedule of the original DDPM paper and the cosine schedule of Nichol and Dhariwal. The second plots $\log_{10}\mathrm{SNR}$, which straightens out the exponential decay and makes the difference in slope the whole story. The dashed line marks the timestep selected above, so you can read off exactly where the demo currently is.

Top: $\bar\alpha_t$ for both schedules. Bottom: $\log_{10}\mathrm{SNR}(t)$. The vertical dashed line is the current timestep.

The linear schedule — $\beta$ from $10^{-4}$ to $0.02$ — spends the first half of its budget barely touching the image and then destroys it quickly near the end. The cosine schedule was designed to fix exactly that imbalance: it removes signal more evenly, so the middle of training is not wasted on timesteps where nothing interesting is happening.

Both curves are computed with accent for linear and accent-2 for cosine, matching the two buttons in the demo.

4

What the schedule choice changes

A schedule is a curriculum for the denoiser

A schedule sets the distribution of difficulty the network sees during training. Timesteps where $\mathrm{SNR}$ is enormous are almost trivial — the answer is nearly visible — and timesteps where it is tiny are nearly impossible, because the image has been fully forgotten. The interesting range is the band around moderate SNR, and a schedule that spends most of its steps there gives the network more useful gradients per unit of compute.

That is the real reason cosine tends to beat linear at the same number of sampling steps. It is not that one is more mathematically correct; both define a valid $q$, and the reverse process can in principle invert either. It is that cosine puts more of its resolution where the denoising problem is hard, so a sampler with a fixed step budget spends its steps wisely.

The choice interacts with everything downstream. More steps at high SNR sharpen detail; more steps near the end fix global composition, because at low SNR the model can only move large-scale structure. Later parts will pick a sampler and a step count, and those choices are usually made by looking at the SNR curve and spending steps where it is changing fastest.

⚠ Do not confuse the training schedule with the sampling step count. Training uses all 1,000 (or more) timesteps because each one is one cheap regression. Sampling uses as few as 10–50 reverse steps, chosen by a different rule. The two numbers are unrelated, and conflating them is the most common misreading of a diffusion paper.

One more quantity is worth naming before we leave: at any $t$, the quantity $\sqrt{1-\bar\alpha_t}$ is the standard deviation of the noise that has been added, and it is exactly what the network will be asked to predict. The denoiser's target is the noise, and the score function is that same prediction up to a known factor — but that is the subject of the fifth part.

5

Where this shows up

Forward processes well beyond images

Probability

Gaussians compose

The closed form is just the fact that a sum of independent Gaussians is Gaussian, and the scaling keeps the result in the same family. The probability guide develops the algebra; the probabilistic ML chapter uses it to derive the variance-preserving forward chain from scratch.

Audio & video

Noise is modality-agnostic

Nothing in $q(x_t \mid x_0)$ cares whether $x_0$ is an image, a spectrogram or a video clip. The same closed form runs on all of them, which is why the later parts of this volume reuse this exact machinery in sound and video.

If you understand that every timestep is one Gaussian sample away from the clean data, you understand the first half of diffusion. The second half is the interesting half: given $x_t$, estimate the noise that was added, and then use that estimate to walk backwards. The next part trains a real denoiser in the browser and watches its loss fall.

Further reading

The forward process appears in every diffusion paper, usually without much ceremony, because it is the easy half. These references give it the care it deserves, and the schedules they introduce are the ones used in production systems today.

Sohl-Dickstein and coauthors introduced the non-equilibrium thermodynamics framing; Ho, Jain and Abbeel turned it into the modern denoising objective; Nichol and Dhariwal fixed the schedule; and Song and colleagues show how far the idea generalises. If you read one section, read the schedule comparison in Nichol and Dhariwal.

Cheat sheet

SymbolMeaning here
$\beta_t$Noise added at step $t$; small, and fixed in advance
$\alpha_t = 1-\beta_t$Signal kept at step $t$
$\bar\alpha_t = \prod_{s\le t}\alpha_s$Cumulative signal after $t$ steps
$q(x_t\mid x_0)$$\mathcal{N}(\sqrt{\bar\alpha_t}\,x_0,\,(1-\bar\alpha_t)I)$ — the closed form
$\sqrt{1-\bar\alpha_t}$Standard deviation of the added noise; the denoiser's target
SNR$(t)$$\bar\alpha_t/(1-\bar\alpha_t)$; how easy denoising is at $t$
Linear schedule$\beta$ ramps $10^{-4}\to0.02$; destroys late and fast
Cosine scheduleRemoves signal evenly; more useful steps in the middle
6

Check your understanding

0/4 answered