Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

A clip is a 3-D signal

One more axis, and an ordering

The strip below is a few frames of a moving pattern, drawn frame by frame. Read it left to right and it is a signal indexed by time as well as by the two spatial axes, and the temporal index carries something the spatial ones do not: a direction. Frame $k$ is the past of frame $k+1$, and a model that predicts the next frame is forbidden from looking ahead. That arrow of time is what makes a video model more than a stack of image models, and every architectural choice that follows is a response to it.

The naive extension is obvious: treat the video as a batch of images and diffuse each frame independently. It fails immediately, because nothing ties the frames together and the result flickers. The useful extension is to let the model see the time axis explicitly, which raises a cost problem before it raises a modelling one. A clip has $F$ times as many pixels as a frame, and full attention over all of them is quadratic in a number that just grew by a factor of $F$.

A strip of frames drawn with GenMedia.raster. The vertical guides group frames into the temporal compression window of the causal VAE.

The pixel count is not the only thing that grows. Latent tokens, attention operations and memory all scale with the temporal axis, and the way they scale is the subject of the rest of this part.

💡 By the end of this part you'll be able to explain why the VAE is causal, why attention is factorised into spatial and temporal passes, and estimate the token count and the attention ratio for a clip of a given length and resolution.
2

The causal VAE

Compress time without peeking ahead

Latent diffusion already compresses space by a factor of eight per axis before the diffusion model ever sees the data. Video does the same and adds a temporal axis, typically compressing four frames into one latent frame. A three-second clip at 24 frames per second is 72 frames; after a temporal factor of four it is 18 latent frames, and the diffusion model works on that. The compression is not cosmetic — it is the difference between a transformer that fits and one that does not.

The autoencoder must be causal along time. A convolutional encoder that looked at both directions would let latent frame $k$ depend on pixel frame $k+1$, which would leak the future into the representation and make autoregressive or streaming use impossible. The fix is causal temporal convolutions: the kernel only sees current and past frames, so latent frame $k$ is a function of frames up to $k$. The property is what lets one model serve both full-sequence diffusion and the frame-by-frame rollout discussed in the next chapter.

⚠ Temporal compression is a prior, not a free lunch. Four frames folded into one latent frame means the decoder must invent the motion between them. Fast motion, cuts and fine temporal texture are exactly what a large compression factor loses, which is why the factor is usually small compared with the spatial one.
3

Patches and attention

Full 3-D against the factorised form

After compression, the latent clip is cut into spatiotemporal patches and flattened into a token sequence, just as an image is. The count is easy to read off: latent frames times the patches per frame. The trouble is attention. Full 3-D attention lets every token see every other token, which is quadratic in that count and, for a clip of any length, astronomically expensive.

The standard fix is to factorise the attention into two passes. A spatial pass runs full attention within each frame, letting positions on a frame talk to each other; a temporal pass runs attention across frames at each position, letting a patch look at the same patch over time. The two passes together see the whole clip, but their cost is the sum of two much smaller quadratics rather than one big one. The bars below compare them at the current settings, on a log scale because the gap is large.

Attention operations on a log scale at the current frames, resolution and patch size. The ratio is printed in the readout.

💡 Factorisation is the single most important trick in video diffusion. It does not just make attention cheaper; it gives the model an explicit spatial module and an explicit temporal module, each of which can be regularised and conditioned separately.
4

The token budget

Why a clip costs about a hundred images

The plot traces the two attention costs as the clip lengthens at the current resolution and patch size. The factorised curve grows, but gently; the full 3-D curve climbs the log axis almost twice as fast, because its cost is quadratic in a token count that is itself linear in frame count. This is the whole reason a video backbone is not an image backbone with a longer sequence: at any useful clip length the full 3-D form is unaffordable, and the factorised form is what makes the problem finite.

The budget is also measured against a single image, which is the comparison a user feels. The number of latent tokens in the clip divided by the tokens in one latent frame is the temporal factor — with a compression of four, a 96-frame clip holds 24 latent frames' worth of tokens. Attention, though, grows faster than tokens, so the honest summary is that a few-second clip costs tens to hundreds of times an image depending on whether attention is full or factorised, and that most of the engineering budget goes into keeping that multiplier down.

Attention operations against clip length, log scale, for full 3-D and factorised attention at the settings above.

Token count grows linearly in frames; full 3-D attention grows quadratically in it. The gap between the curves is the attention ratio.

5

Where this shows up

Every video backbone makes these three choices

Video

Latent video diffusion

Sora, Veo and their contemporaries are all causal 3-D VAEs plus a spatiotemporal transformer, and all of them factorise attention. The latent diffusion machinery is unchanged except that the compressor is now three-dimensional and the compressor's causality is a requirement rather than a nicety.

Backbones

From U-Net to DiT over time

Early video models inflated the U-Net with a temporal axis at the coarse levels; current ones patchify the latent clip and run a DiT whose attention is factorised. The sampler is the same story, with a step count that now multiplies an already larger per-step cost.

The next chapter takes the temporal axis seriously as a modelling problem rather than a cost problem: how to condition on a first frame, how errors accumulate over a long rollout, and why native audio and world models are the frontier claims of the moment.

Further reading

Video generation moved from inflating image models to designing around the temporal axis. These papers mark the arc: the first latent video model, the causal compressor, the factorised transformer, and the scaling studies that made clip length a headline number.

Ho and coauthors factorise the image backbone into spatial and temporal components; Blattmann and coauthors put the video model in a latent space; Blattmann and coauthors add the causal 3-D autoencoder; and Brooks and coauthors report the spatiotemporal transformer at scale.

Cheat sheet

TermMeaning here
Causal 3-D VAECompresses space and time; temporal kernel sees only current and past frames
Temporal compressionTypically 4 pixel frames per latent frame; bounds the token count
Latent frames$\lceil F / t_\text{comp}\rceil$; what the diffusion transformer actually sees
Spatiotemporal patchesLatent clip cut into tokens, exactly as an image is patchified
Full 3-D attentionEvery token attends to every token; quadratic in clip tokens
Factorised attentionSpatial pass within frames plus temporal pass across frames; a sum of two quadratics
Token budgetTokens grow linearly in frames; attention grows quadratically
Cost vs an imageTens to hundreds of times an image, depending on length and attention form
7

Check your understanding

0/4 answered