Latent diffusion
A 512×512 photograph is three quarters of a million numbers, and running a denoiser over all of them at every step is ruinous. Latent diffusion refuses the premise: train an autoencoder that squeezes the image into a small grid of feature values, run the entire diffusion process there, and decode once at the end. The compressor is lossy, so every generation carries a reconstruction cost — and understanding where that cost lands is most of what separates a good latent model from a blurry one.
Why compress
The diffusion cost is quadratic in resolution
The reverse process is a sequence of network evaluations, and the network is mostly convolution and attention over a spatial grid. Convolutions cost roughly linearly in the number of pixels; attention costs quadratically in the number of tokens. Doubling the side length of an image quadruples the pixel count and quadruples attention's bill four times over. At 1024×1024 there are more than three million values per image, and a hundred sampling steps means the network sees all of them a hundred times.
The fix is to stop diffusing pixels. An autoencoder is trained once to map an image to a latent of much smaller spatial extent and back again, and the diffusion model is trained and sampled entirely in that latent space. The decoder is run exactly once, on the final latent, to produce the picture. Most of the compute disappears with the resolution, and the expensive generative model works on a grid of tens of thousands of values rather than hundreds of thousands.
This is not free compression. The autoencoder throws information away, and whatever it throws away can never be recovered by the decoder — no amount of diffusion quality puts back detail the latent never stored. The rest of this part is about how much is thrown away and where it shows.
The encoder and the latent
Downsample, then diffuse the small thing
A latent diffusion encoder is a stack of strided convolutions that reduces the spatial side length by a fixed factor — commonly eight — while increasing the channel count. A 512×512 RGB image becomes a 64×64 grid with four channels, so roughly 16,000 values carry the image instead of 786,000. Those four channels are not colours and do not look like an image; they are learned features, and a good latent has the property that its values are roughly centred and roughly unit variance so that the diffusion process's noise assumptions hold.
The demo below replaces the learned encoder with plain average pooling, which is the crudest possible compressor and makes the cost visible without a training run. The source image is the 64×64 test picture; the middle panel is the latent at the current factor; the third panel is the reconstruction after nearest-neighbour upsampling. The MSE under the panels is the squared error between source and reconstruction, averaged over every channel.
Average pooling then nearest-neighbour upsampling. The latent panel really is the tiny grid the diffusion model would operate on — here shown at pixel scale so you can see how few values remain.
Pooling averages each $f\times f$ block; upsampling repeats it. A trained decoder is far better than this, but the shape of the error is the same.
The reconstruction cost
What pooling throws away, and where a real decoder differs
Averaging a block destroys every variation inside it. Flat regions — skies, walls, skin — lose almost nothing, because the average is what they already look like. High-frequency content loses everything: text becomes illegible, a face's eyes blur, fine texture turns to mush. The chart below measures exactly that, raising the factor and watching the MSE climb.
Reconstruction MSE against the downsample factor, for factors 2, 4, 8, 16 and 32. The current factor is highlighted.
Real VAEs are trained with a mixture of reconstruction loss and a perceptual term, usually features from a pretrained network or an adversarial loss, so the decoder learns to hallucinate plausible texture rather than average it. That is what makes an 8× latent workable: the decoder does not have to recover the exact lost pixels, only a distributionally similar image.
A pure MSE objective pushes the decoder toward blur, because the mean of several plausible textures minimises squared error and looks like nothing.
The failure modes are worth naming because they show up in generated images. Compression is worst for content with high-frequency structure and a strong prior about what it should look like: small text, distant faces, repeated fine patterns such as fabric or foliage. When the latent cannot represent it, the decoder invents it, and it invents the same plausible wrong texture every time. That is the origin of the characteristic latent-diffusion artifacts.
Pixel versus latent budget
Tokens scale with the square of the factor
The quantity that matters downstream is the number of tokens the generative network sees. Downsampling by a factor $f$ on each side reduces the token count by $f^2$: an 8× latent has $1/64$ as many spatial positions as the image. Since attention cost is quadratic in token count and convolutional cost linear, the saving compounds. At $f=8$ attention inside the latent operates on a sequence 64 times shorter, which is 4096 times less work per attention layer than the pixel equivalent.
Spatial token count for a 64×64 image against the downsample factor. Dashed: pixel tokens (constant at 4096). Solid: latent tokens, $(64/f)^2$. The marker is the current factor.
The latent also carries four channels instead of three, so the raw value count is $4/3$ times the spatial token count — a rounding error next to the $f^2$ saving.
The model does not get all of that back as speed, because it is usually trained at the higher capacity the latent budget affords. The trade is deliberate: latent diffusion spends the saved compute on a larger network and more attention, which is why modern latent models are far bigger than the early pixel-space ones. The relevant comparison is not pixels against latents at equal parameters but pixels against latents at equal wall-clock time, and the latent side wins by a wide margin at high resolution.
Where this shows up
Every production image and video model
Stable Diffusion and SDXL
The original latent diffusion paper established the 8×, four-channel VAE, and SDXL kept it while moving the text conditioning to two encoders. The U-Net backbone of those models operates entirely on the 64×64 latent, never on the 512×512 image; the decoder is called once at the very end.
Temporal compression
Video models extend the same idea along the time axis with a causal 3-D VAE, compressing both space and time before the transformer sees anything. The token count then scales as frames, height and width each divided by their own factor — which is the only reason a minute of footage fits in a context window at all.
The same pattern appears in audio and in scientific data: compress to a small learned code, model the code, decode once. Whenever a generative model looks at thousands rather than millions of tokens, an autoencoder is doing the work quietly underneath.
Further reading
The latent-diffusion paper is the reference for the 8× four-channel VAE and for the argument that perceptual compression and semantic generation should be separated. The papers after it are the refinements: better reconstruction objectives, adversarial fine-tuning of the decoder, and the video extension.
Rombach and coauthors introduce the two-stage pipeline; Van den Oord and coauthors give the VQ-VAE origin of learned discrete latents; Esser and coauthors add the first-stage GAN loss that sharpens the decoder; and Blattmann and coauthors carry the compressor into time.
- Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser and Björn Ommer, "High-Resolution Image Synthesis with Latent Diffusion Models", 2022 — the 8× VAE and the two-stage pipeline.
- Aaron van den Oord, Oriol Vinyals and Koray Kavukcuoglu, "Neural Discrete Representation Learning", 2017 — the learned-latent idea in its discrete form.
- Patrick Esser, Robin Rombach and Björn Ommer, "Taming Transformers for High-Resolution Image Synthesis", 2021 — the adversarial reconstruction loss that removes decoder blur.
- Andreas Blattmann and coauthors, "Align your Latents: High-Resolution Video Synthesis with Latent Diffusion Models", 2023 — the spatial then temporal compressor.
Cheat sheet
| Term | Meaning here |
|---|---|
| Encoder | Strided convolutions mapping an image to a small latent grid |
| Latent | Commonly $H/8 \times W/8$ with 4 channels; what the diffusion model sees |
| Decoder | Maps the final latent back to pixels; run exactly once |
| Downsample factor $f$ | Side-length reduction; token count falls by $f^2$ |
| Reconstruction loss | Pixel or perceptual error of the autoencoder; sets the fidelity ceiling |
| Latent scaling | A measured factor applied to latents so the noise model's assumptions hold |
| Artifact source | Information the encoder discarded, hallucinated back by the decoder |
| Compute ratio | Attention work changes by roughly $f^4$ per layer at fixed latent width |