Coherent and long
A clip that is cheap to generate is not the same as a clip that stays coherent as it lengthens. The temporal axis has to be anchored — to a first frame, a keyframe, or a low-resolution sketch — and the choice between generating a whole sequence at once and rolling it out frame by frame decides whether errors stay bounded or accumulate. This part follows drift over a rollout, conditions on a first frame, adds cascades, and ends on the claims about audio and world models.
Conditioning on a first frame
Anchor the sequence before sampling it
Image-to-video generation conditions the model on a single frame and asks it to continue. The conditioning frame is usually handled by concatenating it, un-noised, to the latent clip at every sampling step, so the model can compare what it is generating against a fixed reference. A keyframe model does the same at both ends, or at several interior frames, which turns generation into interpolation between anchors and gives the user a timeline to control.
The strip below shows the two regimes side by side. The top row is a sequence generated with the whole clip in view, so every frame is consistent with every other. The bottom row is an autoregressive rollout, where each frame is produced from the ones before it. Toggle the conditioning switch and frame zero is pinned to the first frame: the rollout stops wandering away from its anchor, which is exactly what I2V conditioning buys.
Top row: full-sequence generation. Bottom row: autoregressive rollout. Each dot is a frame; the vertical line marks the conditioned first frame.
Drift and autoregression
Errors that feed themselves
An autoregressive rollout has a structural problem: each frame is generated from previous generated frames, so any error in frame $k$ becomes part of the input to frame $k+1$. If the error is damped, the rollout stays near the truth; if it is amplified, it compounds and the clip drifts. The plot models that with a decay factor, where each new frame inherits a fraction of the previous drift plus a fresh perturbation. A decay close to one grows without bound; a decay well below one stays flat.
Full-sequence diffusion does not have this failure mode, because every frame is generated with the whole clip in view and no frame is an input to another. It buys coherence by attending to all frames at once, at the cost of the quadratic attention from the previous part. This is the central trade of long video: autoregression is cheap per frame but drifts, and full-sequence generation is coherent but its cost grows with clip length. Anchoring the rollout to a conditioned first frame lowers the effective decay, which is why I2V makes long rollouts stable and why the toggle above flattens the curve below.
Per-frame drift magnitude against frame index. The autoregressive curve follows the decay factor; the full-sequence curve stays bounded.
Drift is generated by Diffusion.temporalDrift, whose decay parameter is the fraction of the previous error carried forward.
Cascades and long clips
Sketch first, refine after
A temporal cascade produces a long clip by generating a low-resolution version first and then refining it with one or more super-resolution stages along the time axis. Because the expensive stages only ever see a short window of the sketch, the cost grows far more slowly than the clip length, and the base pass fixes global motion before detail is spent. This is the same idea as the spatial cascades used for high-resolution images, applied to the temporal axis; the level slider shows the trade, where each added stage lifts the output resolution and raises the total compute gently.
Relative compute per cascade level, with the output resolution in the label. The highlighted bar is the selected level.
The other long-clip strategy is to extend the context window rather than roll out. A model trained on clips of a fixed length can be run with a sliding window and a small number of overlap frames, with the overlapping latents anchored so that consecutive windows agree. The result is neither pure autoregression nor a single full-sequence pass, but a middle ground that keeps a bounded working set while still conditioning on a substantial slice of the past.
Audio and the world-model claim
Native sound, and what the simulator claim means
Sound is the newest axis. Early video models produced silent clips and audio was dubbed on afterwards, which fails whenever a sound has to line up with the event that made it. Veo 3 and its contemporaries generate video and audio jointly, so a door slam lands on the frame that shows the door and speech is lip-synced to the mouth that is speaking. Joint generation is more than a convenience: synchronisation is a constraint the model can satisfy only if it is modelling both signals together, which is why native audio is treated as an architectural feature rather than a post-processing step.
The world-model claim is the boldest statement in the field. A model that predicts the next moments of a scene accurately enough to be useful for planning is behaving like a simulator of the world, and Sora's framing in those terms was explicit. The claim is worth taking seriously and worth bounding. These models have no physics engine and no persistent state; they are trained to produce plausible frames, and plausibility coincides with physical correctness only where the training distribution makes it so. Occlusion, object permanence over long spans, and causal interventions are the documented failures, and a world model whose laws change when the prompt changes is a generator with a useful prior rather than a simulator.
Where this shows up
The frontier video systems
I2V, keyframes and camera control
Every consumer video interface exposes a first-frame image, a keyframe timeline, or a camera path, and each is a conditioning signal on the spatiotemporal latent. Camera control is the structural cousin of ControlNet: a pose or trajectory signal injected the same way, on the temporal axis.
Native sound and worlds
Joint audio-video generation and the world-model framing are the two claims that push video beyond an image model in a loop. Audio conditioning reuses the generative-media toolbox on a new axis, and the simulation claim is the one to hold to a strict standard rather than accept from a demo reel.
Read as a whole, the video chapters are one argument: the temporal axis imposes an ordering, that ordering forces the VAE to be causal and the attention to be factorised, and it raises a coherence problem that conditioning, cascades and the choice of rollout all address. The sampler, the schedule and the backbone from the earlier parts are unchanged underneath.
Further reading
These references cover the two halves of the long-video problem: how to condition and extend a clip, and how far the resulting model can be said to simulate anything.
Ho and coauthors and Blattmann and coauthors establish conditioning and long video; Harvey and coauthors and Yin and coauthors handle long-horizon and streaming generation; and the Sora report is the primary source for the world-model framing.
- Andreas Blattmann, Tim Dockhorn, Sumith Kulal and coauthors, "Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets", 2023 — image-to-video conditioning and the causal 3-D VAE.
- William Harvey, Saeid Naderiparizi, Vaden Masrani, Christian Weilbach and Frank Wood, "Flexible Diffusion Modeling of Long Videos", 2022 — long-horizon generation with a bounded working set.
- Tianwei Yin, Qiang Zhang, Richard Zhang and coauthors, "From Slow Bidirectional to Fast Autoregressive Video Diffusion Models", 2024 — the streaming and autoregressive framing.
- OpenAI, "Video Generation Models as World Simulators", 2024 — the world-model claim, stated and qualified.
- Google DeepMind, "Veo", 2025 — native synchronised audio and video generation in one model.
Cheat sheet
| Term | Meaning here |
|---|---|
| I2V conditioning | Pin a first frame, un-noised, at every sampling step |
| Keyframe conditioning | Anchor several frames; generation becomes interpolation |
| Autoregressive rollout | Each frame is generated from previous frames; errors compound |
| Full-sequence diffusion | All frames generated at once; coherent, with quadratic attention |
| Drift | Accumulated error over a rollout; Diffusion.temporalDrift |
| Temporal cascade | Low-resolution sketch plus super-resolution stages along time |
| Sliding window | Bounded context with anchored overlap frames |
| Native audio | Video and sound generated jointly so events stay synchronised |
| World-model claim | A strong learned prior over scene evolution, not a physics engine |