Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

The resolution pyramid

Shrink the grid, widen the channels

A pyramid is the natural shape for a denoiser because the two quantities move in opposite directions. Early layers need to resolve fine spatial detail but can afford few channels; deep layers have lost most of the spatial resolution and compensate with many channels, so each position summarises a large region of the image. Three downsample stages take a 64×64 latent to 8×8 while the channel count climbs from 320 to 1280, and the decoder runs the same schedule backwards.

The schematic below is drawn to scale: the width of each block is proportional to its channel count, and the two columns are the encoder and decoder. Change the base channel count and the blocks per level and watch the U widen, then read the parameter totals underneath. The counts are an envelope from Diffusion.unetBreakdown — a realistic order of magnitude rather than a specific checkpoint.

A four-level U-Net. Dashed horizontal arrows are skip connections from encoder to decoder; the curved connector at the bottom is the bottleneck. Block width scales with channels.

💡 By the end of this part you'll be able to read a U-Net diagram as a table of resolutions and channels, say what each skip connection carries, explain why self-attention lives at the bottom of the U, and estimate where the parameters actually sit.
2

Skips and why they matter

The wire that keeps the detail

Downsampling is lossy. By the time the signal reaches the bottleneck it is an 8×8 grid with 1280 channels, which is a fine summary of global structure and a hopeless record of where an edge was. A plain autoencoder would have to regenerate the missing detail from that summary. The U-Net's trick is to bypass the loss: at every level, the encoder's feature map is copied across and concatenated onto the decoder's feature map at the same resolution. The decoder therefore has both the coarse content coming up from below and the sharp, local features coming from the left.

This is why the denoiser can be so good at the fine-grained part of the job. At high signal-to-noise the answer is mostly a small cleanup of the input, and the skip connections hand the decoder exactly the input's detail to clean. At low signal-to-noise the skips are nearly noise and the network leans on the bottleneck's global view instead. The same wires serve both regimes, which is part of why one architecture covers the whole trajectory.

⚠ Skips are not free. Concatenating an encoder map onto a decoder map doubles the decoder's input channels at that level, so the first convolution after each skip is the expensive one. The wider blocks in the schematic's right column are that cost made visible.
3

Where attention sits

Convolution sees locally; attention sees the whole grid

A convolution mixes information only within its kernel, so a stack of them has an effective receptive field that grows slowly with depth. Self-attention has no such limit: every position attends to every other position in one layer, which is exactly what is needed to keep a face consistent with a body or a horizon straight across the frame. The catch is cost. Attention is quadratic in the number of positions, so it is ruinous at 64×64 and cheap at 8×8.

The U-Net resolves the tension by placing self-attention only at the coarse levels, where the sequence is short enough that the quadratic term stays small. A typical layout puts a self-attention block after the residual blocks at the two deepest resolutions, which is why the schematic tags those levels. Everything above stays purely convolutional, spending its budget on local detail at a resolution where attention could not fit anyway.

This is the design decision that the DiT architecture eventually inverts. A U-Net keeps the pyramid and sprinkles attention where it fits; a transformer runs attention everywhere and replaces the pyramid with a single-resolution sequence of patches. The next part makes that comparison concrete.

4

Conditioning and time

Two extra inputs: which step, and which prompt

A denoiser is asked to do different jobs at different times. At $t$ near zero it removes a whisper of noise; at large $t$ it must invent structure from almost nothing. The network is told which regime it is in by a timestep embedding: a fixed sinusoidal encoding of $t$ is projected through a small MLP and then injected into every residual block, usually by predicting a scale and a shift that modulate the block's normalisation. So the same weights behave very differently at different noise levels, cheaply.

Text conditioning arrives by a second mechanism. Prompt tokens are encoded once by a text encoder, and then each spatial position of the U-Net attends to those tokens with cross-attention: queries come from the image features and keys and values come from the text. A patch representing fur can attend to the word "cat" and pull in the relevant direction, which is how a class or caption steers generation without retraining the denoiser. Cross-attention is inserted at the same coarse levels as self-attention, for the same cost reason.

Two knobs therefore coexist inside the network: a scalar clock that says how much noise remains, and a set of token features that say what to draw. Keeping them separate is what lets one model serve a thousand noise levels and an unbounded space of prompts.

5

Counting parameters

Where the weights actually live

A residual block is two $3\times3$ convolutions, and a convolution from $c_\text{in}$ to $c_\text{out}$ channels costs $c_\text{in}\cdot c_\text{out}\cdot 9$ weights. That product is why the deep levels dominate: the spatial grid shrinks by four at each downsample, but the channel count doubles, so the parameter count of a block grows even as its feature map shrinks. Attention adds a further $2c^2$ per block at the levels that carry it.

Parameters per resolution level from Diffusion.unetBreakdown, at the base channels and blocks-per-level set by the sliders above.

The bars are the same numbers the readout prints, drawn at the same scale as the schematic: the wide blocks at the bottom of the U are the expensive ones. Increase the blocks slider and every level grows together; increase the base channels and the deep levels grow fastest, because the channel counts there are multiples of the base.

Doubling the base channels roughly quadruples most of the parameter count, which is why width is a blunt instrument and depth a cheaper one.

6

Where this shows up

A decade of image and video denoisers

Image

Stable Diffusion and Imagen

The latent diffusion U-Net and the pixel-space Imagen stack are both this architecture with different widths. Dhariwal and Nichol's ablation is the reference for where attention and how many residual blocks to place.

Video & audio

Inflated and 1-D U-Nets

Video U-Nets inflate every 2-D block into a spatio-temporal one, adding a temporal axis only at the coarse levels to keep the cost down. Audio models flatten a spectrogram into a 1-D sequence and run a U-Net along time with the same skip structure.

The U-Net is the workhorse that made diffusion practical, and its inductive biases — locality, translation equivariance, a resolution pyramid — are exactly what the DiT gives up in exchange for scale. Which of those two bets wins is the subject of the next part.

Further reading

The U-Net predates diffusion by years and was borrowed almost unchanged. The two papers that matter most for the diffusion variant are the DDPM architecture and the Dhariwal–Nichol ablation, which together fix where attention goes and how many blocks each level carries.

Ronneberger and coauthors introduce the shape for biomedical segmentation; Ho and coauthors bring it into diffusion; Dhariwal and Nichol tune its width, depth and attention placement against a GAN baseline; and Rombach and coauthors move it into the latent space.

Cheat sheet

TermMeaning here
Resolution pyramidGrid shrinks by 2 per stage; channels double in compensation
Skip connectionConcatenates an encoder map onto the decoder at the same resolution
Residual blockTwo $3\times3$ convs; $c_\text{in}c_\text{out}\cdot 9$ weights per conv
Self-attentionEvery position sees every position; placed at the coarse levels only
Cross-attentionImage positions attend to prompt tokens; how text steers generation
Timestep embeddingSinusoidal encoding of $t$ projected and injected into every block
BottleneckLowest resolution, most channels; holds the global summary
Parameter massConcentrated at the deep levels, where channels are widest
8

Check your understanding

0/4 answered