The U-Net backbone
The denoiser for a latent diffusion model is a U-Net: an encoder that shrinks the spatial grid while widening the channels, a bottleneck, and a decoder that reverses the trip. Skip connections carry the fine detail across the middle so the decoder is not forced to reinvent it, and attention is bolted in at the coarse levels where a global view is affordable. This part draws the architecture to scale and lets you change its width and depth, with the parameter bill updating as you do.
The resolution pyramid
Shrink the grid, widen the channels
A pyramid is the natural shape for a denoiser because the two quantities move in opposite directions. Early layers need to resolve fine spatial detail but can afford few channels; deep layers have lost most of the spatial resolution and compensate with many channels, so each position summarises a large region of the image. Three downsample stages take a 64×64 latent to 8×8 while the channel count climbs from 320 to 1280, and the decoder runs the same schedule backwards.
The schematic below is drawn to scale: the width of each block is proportional to its channel count, and the two columns are the encoder and decoder. Change the base channel count and the blocks per level and watch the U widen, then read the parameter totals underneath. The counts are an envelope from Diffusion.unetBreakdown — a realistic order of magnitude rather than a specific checkpoint.
A four-level U-Net. Dashed horizontal arrows are skip connections from encoder to decoder; the curved connector at the bottom is the bottleneck. Block width scales with channels.
Skips and why they matter
The wire that keeps the detail
Downsampling is lossy. By the time the signal reaches the bottleneck it is an 8×8 grid with 1280 channels, which is a fine summary of global structure and a hopeless record of where an edge was. A plain autoencoder would have to regenerate the missing detail from that summary. The U-Net's trick is to bypass the loss: at every level, the encoder's feature map is copied across and concatenated onto the decoder's feature map at the same resolution. The decoder therefore has both the coarse content coming up from below and the sharp, local features coming from the left.
This is why the denoiser can be so good at the fine-grained part of the job. At high signal-to-noise the answer is mostly a small cleanup of the input, and the skip connections hand the decoder exactly the input's detail to clean. At low signal-to-noise the skips are nearly noise and the network leans on the bottleneck's global view instead. The same wires serve both regimes, which is part of why one architecture covers the whole trajectory.
Where attention sits
Convolution sees locally; attention sees the whole grid
A convolution mixes information only within its kernel, so a stack of them has an effective receptive field that grows slowly with depth. Self-attention has no such limit: every position attends to every other position in one layer, which is exactly what is needed to keep a face consistent with a body or a horizon straight across the frame. The catch is cost. Attention is quadratic in the number of positions, so it is ruinous at 64×64 and cheap at 8×8.
The U-Net resolves the tension by placing self-attention only at the coarse levels, where the sequence is short enough that the quadratic term stays small. A typical layout puts a self-attention block after the residual blocks at the two deepest resolutions, which is why the schematic tags those levels. Everything above stays purely convolutional, spending its budget on local detail at a resolution where attention could not fit anyway.
This is the design decision that the DiT architecture eventually inverts. A U-Net keeps the pyramid and sprinkles attention where it fits; a transformer runs attention everywhere and replaces the pyramid with a single-resolution sequence of patches. The next part makes that comparison concrete.
Conditioning and time
Two extra inputs: which step, and which prompt
A denoiser is asked to do different jobs at different times. At $t$ near zero it removes a whisper of noise; at large $t$ it must invent structure from almost nothing. The network is told which regime it is in by a timestep embedding: a fixed sinusoidal encoding of $t$ is projected through a small MLP and then injected into every residual block, usually by predicting a scale and a shift that modulate the block's normalisation. So the same weights behave very differently at different noise levels, cheaply.
Text conditioning arrives by a second mechanism. Prompt tokens are encoded once by a text encoder, and then each spatial position of the U-Net attends to those tokens with cross-attention: queries come from the image features and keys and values come from the text. A patch representing fur can attend to the word "cat" and pull in the relevant direction, which is how a class or caption steers generation without retraining the denoiser. Cross-attention is inserted at the same coarse levels as self-attention, for the same cost reason.
Two knobs therefore coexist inside the network: a scalar clock that says how much noise remains, and a set of token features that say what to draw. Keeping them separate is what lets one model serve a thousand noise levels and an unbounded space of prompts.
Counting parameters
Where the weights actually live
A residual block is two $3\times3$ convolutions, and a convolution from $c_\text{in}$ to $c_\text{out}$ channels costs $c_\text{in}\cdot c_\text{out}\cdot 9$ weights. That product is why the deep levels dominate: the spatial grid shrinks by four at each downsample, but the channel count doubles, so the parameter count of a block grows even as its feature map shrinks. Attention adds a further $2c^2$ per block at the levels that carry it.
Parameters per resolution level from Diffusion.unetBreakdown, at the base channels and blocks-per-level set by the sliders above.
The bars are the same numbers the readout prints, drawn at the same scale as the schematic: the wide blocks at the bottom of the U are the expensive ones. Increase the blocks slider and every level grows together; increase the base channels and the deep levels grow fastest, because the channel counts there are multiples of the base.
Doubling the base channels roughly quadruples most of the parameter count, which is why width is a blunt instrument and depth a cheaper one.
Where this shows up
A decade of image and video denoisers
Stable Diffusion and Imagen
The latent diffusion U-Net and the pixel-space Imagen stack are both this architecture with different widths. Dhariwal and Nichol's ablation is the reference for where attention and how many residual blocks to place.
Inflated and 1-D U-Nets
Video U-Nets inflate every 2-D block into a spatio-temporal one, adding a temporal axis only at the coarse levels to keep the cost down. Audio models flatten a spectrogram into a 1-D sequence and run a U-Net along time with the same skip structure.
The U-Net is the workhorse that made diffusion practical, and its inductive biases — locality, translation equivariance, a resolution pyramid — are exactly what the DiT gives up in exchange for scale. Which of those two bets wins is the subject of the next part.
Further reading
The U-Net predates diffusion by years and was borrowed almost unchanged. The two papers that matter most for the diffusion variant are the DDPM architecture and the Dhariwal–Nichol ablation, which together fix where attention goes and how many blocks each level carries.
Ronneberger and coauthors introduce the shape for biomedical segmentation; Ho and coauthors bring it into diffusion; Dhariwal and Nichol tune its width, depth and attention placement against a GAN baseline; and Rombach and coauthors move it into the latent space.
- Olaf Ronneberger, Philipp Fischer and Thomas Brox, "U-Net: Convolutional Networks for Biomedical Image Segmentation", 2015 — the original encoder–decoder with skips.
- Jonathan Ho, Ajay Jain and Pieter Abbeel, "Denoising Diffusion Probabilistic Models", 2020 — the U-Net as a noise predictor.
- Prafulla Dhariwal and Alex Nichol, "Diffusion Models Beat GANs on Image Synthesis", 2021 — the architecture ablation: width, depth and attention at coarse resolutions.
- Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser and Björn Ommer, "High-Resolution Image Synthesis with Latent Diffusion Models", 2022 — cross-attention for text conditioning.
Cheat sheet
| Term | Meaning here |
|---|---|
| Resolution pyramid | Grid shrinks by 2 per stage; channels double in compensation |
| Skip connection | Concatenates an encoder map onto the decoder at the same resolution |
| Residual block | Two $3\times3$ convs; $c_\text{in}c_\text{out}\cdot 9$ weights per conv |
| Self-attention | Every position sees every position; placed at the coarse levels only |
| Cross-attention | Image positions attend to prompt tokens; how text steers generation |
| Timestep embedding | Sinusoidal encoding of $t$ projected and injected into every block |
| Bottleneck | Lowest resolution, most channels; holds the global summary |
| Parameter mass | Concentrated at the deep levels, where channels are widest |