Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

Conditioning on a first frame

Anchor the sequence before sampling it

Image-to-video generation conditions the model on a single frame and asks it to continue. The conditioning frame is usually handled by concatenating it, un-noised, to the latent clip at every sampling step, so the model can compare what it is generating against a fixed reference. A keyframe model does the same at both ends, or at several interior frames, which turns generation into interpolation between anchors and gives the user a timeline to control.

The strip below shows the two regimes side by side. The top row is a sequence generated with the whole clip in view, so every frame is consistent with every other. The bottom row is an autoregressive rollout, where each frame is produced from the ones before it. Toggle the conditioning switch and frame zero is pinned to the first frame: the rollout stops wandering away from its anchor, which is exactly what I2V conditioning buys.

Top row: full-sequence generation. Bottom row: autoregressive rollout. Each dot is a frame; the vertical line marks the conditioned first frame.

💡 By the end of this part you'll be able to say why conditioning bounds drift, why autoregressive and full-sequence diffusion fail differently, what a temporal cascade adds, and what the world-model claim does and does not assert.
2

Drift and autoregression

Errors that feed themselves

An autoregressive rollout has a structural problem: each frame is generated from previous generated frames, so any error in frame $k$ becomes part of the input to frame $k+1$. If the error is damped, the rollout stays near the truth; if it is amplified, it compounds and the clip drifts. The plot models that with a decay factor, where each new frame inherits a fraction of the previous drift plus a fresh perturbation. A decay close to one grows without bound; a decay well below one stays flat.

Full-sequence diffusion does not have this failure mode, because every frame is generated with the whole clip in view and no frame is an input to another. It buys coherence by attending to all frames at once, at the cost of the quadratic attention from the previous part. This is the central trade of long video: autoregression is cheap per frame but drifts, and full-sequence generation is coherent but its cost grows with clip length. Anchoring the rollout to a conditioned first frame lowers the effective decay, which is why I2V makes long rollouts stable and why the toggle above flattens the curve below.

Per-frame drift magnitude against frame index. The autoregressive curve follows the decay factor; the full-sequence curve stays bounded.

Drift is generated by Diffusion.temporalDrift, whose decay parameter is the fraction of the previous error carried forward.

⚠ A bounded curve is not a correct clip. Low drift means the rollout stays near its anchor, not that it is right. A model can be stable and consistently wrong, and a degraded rollout often looks fine frame by frame while failing the moment it is played as motion. Stability and faithfulness are separate properties.
3

Cascades and long clips

Sketch first, refine after

A temporal cascade produces a long clip by generating a low-resolution version first and then refining it with one or more super-resolution stages along the time axis. Because the expensive stages only ever see a short window of the sketch, the cost grows far more slowly than the clip length, and the base pass fixes global motion before detail is spent. This is the same idea as the spatial cascades used for high-resolution images, applied to the temporal axis; the level slider shows the trade, where each added stage lifts the output resolution and raises the total compute gently.

Relative compute per cascade level, with the output resolution in the label. The highlighted bar is the selected level.

The other long-clip strategy is to extend the context window rather than roll out. A model trained on clips of a fixed length can be run with a sliding window and a small number of overlap frames, with the overlapping latents anchored so that consecutive windows agree. The result is neither pure autoregression nor a single full-sequence pass, but a middle ground that keeps a bounded working set while still conditioning on a substantial slice of the past.

4

Audio and the world-model claim

Native sound, and what the simulator claim means

Sound is the newest axis. Early video models produced silent clips and audio was dubbed on afterwards, which fails whenever a sound has to line up with the event that made it. Veo 3 and its contemporaries generate video and audio jointly, so a door slam lands on the frame that shows the door and speech is lip-synced to the mouth that is speaking. Joint generation is more than a convenience: synchronisation is a constraint the model can satisfy only if it is modelling both signals together, which is why native audio is treated as an architectural feature rather than a post-processing step.

The world-model claim is the boldest statement in the field. A model that predicts the next moments of a scene accurately enough to be useful for planning is behaving like a simulator of the world, and Sora's framing in those terms was explicit. The claim is worth taking seriously and worth bounding. These models have no physics engine and no persistent state; they are trained to produce plausible frames, and plausibility coincides with physical correctness only where the training distribution makes it so. Occlusion, object permanence over long spans, and causal interventions are the documented failures, and a world model whose laws change when the prompt changes is a generator with a useful prior rather than a simulator.

💡 The honest summary is that video models learn a strong prior over how scenes evolve, and that the prior is good enough for short, in-distribution clips. The world-model claim is a research programme about extending that, not a property that follows from sample quality.
5

Where this shows up

The frontier video systems

Conditioning

I2V, keyframes and camera control

Every consumer video interface exposes a first-frame image, a keyframe timeline, or a camera path, and each is a conditioning signal on the spatiotemporal latent. Camera control is the structural cousin of ControlNet: a pose or trajectory signal injected the same way, on the temporal axis.

Audio & simulation

Native sound and worlds

Joint audio-video generation and the world-model framing are the two claims that push video beyond an image model in a loop. Audio conditioning reuses the generative-media toolbox on a new axis, and the simulation claim is the one to hold to a strict standard rather than accept from a demo reel.

Read as a whole, the video chapters are one argument: the temporal axis imposes an ordering, that ordering forces the VAE to be causal and the attention to be factorised, and it raises a coherence problem that conditioning, cascades and the choice of rollout all address. The sampler, the schedule and the backbone from the earlier parts are unchanged underneath.

Further reading

These references cover the two halves of the long-video problem: how to condition and extend a clip, and how far the resulting model can be said to simulate anything.

Ho and coauthors and Blattmann and coauthors establish conditioning and long video; Harvey and coauthors and Yin and coauthors handle long-horizon and streaming generation; and the Sora report is the primary source for the world-model framing.

Cheat sheet

TermMeaning here
I2V conditioningPin a first frame, un-noised, at every sampling step
Keyframe conditioningAnchor several frames; generation becomes interpolation
Autoregressive rolloutEach frame is generated from previous frames; errors compound
Full-sequence diffusionAll frames generated at once; coherent, with quadratic attention
DriftAccumulated error over a rollout; Diffusion.temporalDrift
Temporal cascadeLow-resolution sketch plus super-resolution stages along time
Sliding windowBounded context with anchored overlap frames
Native audioVideo and sound generated jointly so events stay synchronised
World-model claimA strong learned prior over scene evolution, not a physics engine
7

Check your understanding

0/4 answered