Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

Why compress

The diffusion cost is quadratic in resolution

The reverse process is a sequence of network evaluations, and the network is mostly convolution and attention over a spatial grid. Convolutions cost roughly linearly in the number of pixels; attention costs quadratically in the number of tokens. Doubling the side length of an image quadruples the pixel count and quadruples attention's bill four times over. At 1024×1024 there are more than three million values per image, and a hundred sampling steps means the network sees all of them a hundred times.

The fix is to stop diffusing pixels. An autoencoder is trained once to map an image to a latent of much smaller spatial extent and back again, and the diffusion model is trained and sampled entirely in that latent space. The decoder is run exactly once, on the final latent, to produce the picture. Most of the compute disappears with the resolution, and the expensive generative model works on a grid of tens of thousands of values rather than hundreds of thousands.

This is not free compression. The autoencoder throws information away, and whatever it throws away can never be recovered by the decoder — no amount of diffusion quality puts back detail the latent never stored. The rest of this part is about how much is thrown away and where it shows.

💡 By the end of this part you'll be able to say what an 8× downsample does to the token count, explain why the latent has its own channel dimension, predict which image content suffers most from aggressive compression, and quote the compute ratio between pixel and latent diffusion.
2

The encoder and the latent

Downsample, then diffuse the small thing

A latent diffusion encoder is a stack of strided convolutions that reduces the spatial side length by a fixed factor — commonly eight — while increasing the channel count. A 512×512 RGB image becomes a 64×64 grid with four channels, so roughly 16,000 values carry the image instead of 786,000. Those four channels are not colours and do not look like an image; they are learned features, and a good latent has the property that its values are roughly centred and roughly unit variance so that the diffusion process's noise assumptions hold.

The demo below replaces the learned encoder with plain average pooling, which is the crudest possible compressor and makes the cost visible without a training run. The source image is the 64×64 test picture; the middle panel is the latent at the current factor; the third panel is the reconstruction after nearest-neighbour upsampling. The MSE under the panels is the squared error between source and reconstruction, averaged over every channel.

source 64×64×3
latent
reconstruction

Average pooling then nearest-neighbour upsampling. The latent panel really is the tiny grid the diffusion model would operate on — here shown at pixel scale so you can see how few values remain.

Pooling averages each $f\times f$ block; upsampling repeats it. A trained decoder is far better than this, but the shape of the error is the same.

⚠ The latent is not an image. Four channels is a compromise: fewer and the decoder cannot recover texture, more and the diffusion model's attention gets expensive again. Real latents are unbounded floats, so the pipeline scales them by a measured factor before adding noise, and getting that factor wrong is a classic cause of washed-out samples.
3

The reconstruction cost

What pooling throws away, and where a real decoder differs

Averaging a block destroys every variation inside it. Flat regions — skies, walls, skin — lose almost nothing, because the average is what they already look like. High-frequency content loses everything: text becomes illegible, a face's eyes blur, fine texture turns to mush. The chart below measures exactly that, raising the factor and watching the MSE climb.

Reconstruction MSE against the downsample factor, for factors 2, 4, 8, 16 and 32. The current factor is highlighted.

Real VAEs are trained with a mixture of reconstruction loss and a perceptual term, usually features from a pretrained network or an adversarial loss, so the decoder learns to hallucinate plausible texture rather than average it. That is what makes an 8× latent workable: the decoder does not have to recover the exact lost pixels, only a distributionally similar image.

A pure MSE objective pushes the decoder toward blur, because the mean of several plausible textures minimises squared error and looks like nothing.

The failure modes are worth naming because they show up in generated images. Compression is worst for content with high-frequency structure and a strong prior about what it should look like: small text, distant faces, repeated fine patterns such as fabric or foliage. When the latent cannot represent it, the decoder invents it, and it invents the same plausible wrong texture every time. That is the origin of the characteristic latent-diffusion artifacts.

4

Pixel versus latent budget

Tokens scale with the square of the factor

The quantity that matters downstream is the number of tokens the generative network sees. Downsampling by a factor $f$ on each side reduces the token count by $f^2$: an 8× latent has $1/64$ as many spatial positions as the image. Since attention cost is quadratic in token count and convolutional cost linear, the saving compounds. At $f=8$ attention inside the latent operates on a sequence 64 times shorter, which is 4096 times less work per attention layer than the pixel equivalent.

Spatial token count for a 64×64 image against the downsample factor. Dashed: pixel tokens (constant at 4096). Solid: latent tokens, $(64/f)^2$. The marker is the current factor.

The latent also carries four channels instead of three, so the raw value count is $4/3$ times the spatial token count — a rounding error next to the $f^2$ saving.

The model does not get all of that back as speed, because it is usually trained at the higher capacity the latent budget affords. The trade is deliberate: latent diffusion spends the saved compute on a larger network and more attention, which is why modern latent models are far bigger than the early pixel-space ones. The relevant comparison is not pixels against latents at equal parameters but pixels against latents at equal wall-clock time, and the latent side wins by a wide margin at high resolution.

⚠ Downsampling also changes the learning problem. The latent distribution is smoother and more Gaussian than pixels, which is easier for the diffusion model, but the decoder must then map a coarse, noisy latent back to sharp detail. Where the encoder is too aggressive, the decoder's reconstruction error dominates the sampler's error and no sampler improvement is visible.
5

Where this shows up

Every production image and video model

Image

Stable Diffusion and SDXL

The original latent diffusion paper established the 8×, four-channel VAE, and SDXL kept it while moving the text conditioning to two encoders. The U-Net backbone of those models operates entirely on the 64×64 latent, never on the 512×512 image; the decoder is called once at the very end.

Video

Temporal compression

Video models extend the same idea along the time axis with a causal 3-D VAE, compressing both space and time before the transformer sees anything. The token count then scales as frames, height and width each divided by their own factor — which is the only reason a minute of footage fits in a context window at all.

The same pattern appears in audio and in scientific data: compress to a small learned code, model the code, decode once. Whenever a generative model looks at thousands rather than millions of tokens, an autoencoder is doing the work quietly underneath.

Further reading

The latent-diffusion paper is the reference for the 8× four-channel VAE and for the argument that perceptual compression and semantic generation should be separated. The papers after it are the refinements: better reconstruction objectives, adversarial fine-tuning of the decoder, and the video extension.

Rombach and coauthors introduce the two-stage pipeline; Van den Oord and coauthors give the VQ-VAE origin of learned discrete latents; Esser and coauthors add the first-stage GAN loss that sharpens the decoder; and Blattmann and coauthors carry the compressor into time.

Cheat sheet

TermMeaning here
EncoderStrided convolutions mapping an image to a small latent grid
LatentCommonly $H/8 \times W/8$ with 4 channels; what the diffusion model sees
DecoderMaps the final latent back to pixels; run exactly once
Downsample factor $f$Side-length reduction; token count falls by $f^2$
Reconstruction lossPixel or perceptual error of the autoencoder; sets the fidelity ceiling
Latent scalingA measured factor applied to latents so the noise model's assumptions hold
Artifact sourceInformation the encoder discarded, hallucinated back by the decoder
Compute ratioAttention work changes by roughly $f^4$ per layer at fixed latent width
7

Check your understanding

0/4 answered