Video as a 3D signal
An image is a grid with two axes. A video is a grid with three, and the third axis is not like the others: it is ordered, and every frame must be consistent with the ones before it. Video diffusion is mostly the story of making that third axis affordable. A causal 3-D autoencoder compresses it, spatiotemporal patches tokenise it, and factorised attention splits the quadratic cost into a spatial term and a temporal one. This part counts what a clip actually costs.
A clip is a 3-D signal
One more axis, and an ordering
The strip below is a few frames of a moving pattern, drawn frame by frame. Read it left to right and it is a signal indexed by time as well as by the two spatial axes, and the temporal index carries something the spatial ones do not: a direction. Frame $k$ is the past of frame $k+1$, and a model that predicts the next frame is forbidden from looking ahead. That arrow of time is what makes a video model more than a stack of image models, and every architectural choice that follows is a response to it.
The naive extension is obvious: treat the video as a batch of images and diffuse each frame independently. It fails immediately, because nothing ties the frames together and the result flickers. The useful extension is to let the model see the time axis explicitly, which raises a cost problem before it raises a modelling one. A clip has $F$ times as many pixels as a frame, and full attention over all of them is quadratic in a number that just grew by a factor of $F$.
A strip of frames drawn with GenMedia.raster. The vertical guides group frames into the temporal compression window of the causal VAE.
The pixel count is not the only thing that grows. Latent tokens, attention operations and memory all scale with the temporal axis, and the way they scale is the subject of the rest of this part.
The causal VAE
Compress time without peeking ahead
Latent diffusion already compresses space by a factor of eight per axis before the diffusion model ever sees the data. Video does the same and adds a temporal axis, typically compressing four frames into one latent frame. A three-second clip at 24 frames per second is 72 frames; after a temporal factor of four it is 18 latent frames, and the diffusion model works on that. The compression is not cosmetic — it is the difference between a transformer that fits and one that does not.
The autoencoder must be causal along time. A convolutional encoder that looked at both directions would let latent frame $k$ depend on pixel frame $k+1$, which would leak the future into the representation and make autoregressive or streaming use impossible. The fix is causal temporal convolutions: the kernel only sees current and past frames, so latent frame $k$ is a function of frames up to $k$. The property is what lets one model serve both full-sequence diffusion and the frame-by-frame rollout discussed in the next chapter.
Patches and attention
Full 3-D against the factorised form
After compression, the latent clip is cut into spatiotemporal patches and flattened into a token sequence, just as an image is. The count is easy to read off: latent frames times the patches per frame. The trouble is attention. Full 3-D attention lets every token see every other token, which is quadratic in that count and, for a clip of any length, astronomically expensive.
The standard fix is to factorise the attention into two passes. A spatial pass runs full attention within each frame, letting positions on a frame talk to each other; a temporal pass runs attention across frames at each position, letting a patch look at the same patch over time. The two passes together see the whole clip, but their cost is the sum of two much smaller quadratics rather than one big one. The bars below compare them at the current settings, on a log scale because the gap is large.
Attention operations on a log scale at the current frames, resolution and patch size. The ratio is printed in the readout.
The token budget
Why a clip costs about a hundred images
The plot traces the two attention costs as the clip lengthens at the current resolution and patch size. The factorised curve grows, but gently; the full 3-D curve climbs the log axis almost twice as fast, because its cost is quadratic in a token count that is itself linear in frame count. This is the whole reason a video backbone is not an image backbone with a longer sequence: at any useful clip length the full 3-D form is unaffordable, and the factorised form is what makes the problem finite.
The budget is also measured against a single image, which is the comparison a user feels. The number of latent tokens in the clip divided by the tokens in one latent frame is the temporal factor — with a compression of four, a 96-frame clip holds 24 latent frames' worth of tokens. Attention, though, grows faster than tokens, so the honest summary is that a few-second clip costs tens to hundreds of times an image depending on whether attention is full or factorised, and that most of the engineering budget goes into keeping that multiplier down.
Attention operations against clip length, log scale, for full 3-D and factorised attention at the settings above.
Token count grows linearly in frames; full 3-D attention grows quadratically in it. The gap between the curves is the attention ratio.
Where this shows up
Every video backbone makes these three choices
Latent video diffusion
Sora, Veo and their contemporaries are all causal 3-D VAEs plus a spatiotemporal transformer, and all of them factorise attention. The latent diffusion machinery is unchanged except that the compressor is now three-dimensional and the compressor's causality is a requirement rather than a nicety.
From U-Net to DiT over time
Early video models inflated the U-Net with a temporal axis at the coarse levels; current ones patchify the latent clip and run a DiT whose attention is factorised. The sampler is the same story, with a step count that now multiplies an already larger per-step cost.
The next chapter takes the temporal axis seriously as a modelling problem rather than a cost problem: how to condition on a first frame, how errors accumulate over a long rollout, and why native audio and world models are the frontier claims of the moment.
Further reading
Video generation moved from inflating image models to designing around the temporal axis. These papers mark the arc: the first latent video model, the causal compressor, the factorised transformer, and the scaling studies that made clip length a headline number.
Ho and coauthors factorise the image backbone into spatial and temporal components; Blattmann and coauthors put the video model in a latent space; Blattmann and coauthors add the causal 3-D autoencoder; and Brooks and coauthors report the spatiotemporal transformer at scale.
- Jonathan Ho, Tim Salimans, Alexey Gritsenko and coauthors, "Video Diffusion Models", 2022 — the factorised spatial and temporal attention introduced for video.
- Andreas Blattmann, Robin Rombach, Huan Ling and coauthors, "Align your Latents: High-Resolution Video Synthesis with Latent Diffusion Models", 2023 — video diffusion in a compressed latent.
- Andreas Blattmann, Tim Dockhorn, Sumith Kulal and coauthors, "Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets", 2023 — the causal 3-D VAE and data curation.
- Tim Brooks, Bill Peebles, Connor Holmes and coauthors, "Video Generation Models as World Simulators", 2024 — spatiotemporal patches and variable resolution and duration.
Cheat sheet
| Term | Meaning here |
|---|---|
| Causal 3-D VAE | Compresses space and time; temporal kernel sees only current and past frames |
| Temporal compression | Typically 4 pixel frames per latent frame; bounds the token count |
| Latent frames | $\lceil F / t_\text{comp}\rceil$; what the diffusion transformer actually sees |
| Spatiotemporal patches | Latent clip cut into tokens, exactly as an image is patchified |
| Full 3-D attention | Every token attends to every token; quadratic in clip tokens |
| Factorised attention | Spatial pass within frames plus temporal pass across frames; a sum of two quadratics |
| Token budget | Tokens grow linearly in frames; attention grows quadratically |
| Cost vs an image | Tens to hundreds of times an image, depending on length and attention form |