Video in, long context
A still image is one grid of patches — patches being the equal square tiles an image is cut into, each of which the model treats as one token, the smallest unit it reads and writes. A video is that grid repeated in time, and time does not compress for free: thirty seconds at 24 frames per second is 720 pictures, where a frame is one still image from the sequence and fps, short for frames per second, is how many of them each second of video contains. If each of those pictures costs hundreds of tokens the sequence is longer than most context windows were ever built for. The context window is the fixed maximum number of tokens the model can hold in front of itself at once, the fixed-size suitcase every request has to be packed into, and a video very quickly becomes too bulky for the case. Every video-capable model is therefore a story about sampling — which frames to keep, how to encode their positions, what to hold in memory while the user is still streaming, meaning sending the video in as it arrives rather than all at once, and how large the key-value cache grows behind it all. The KV cache is easiest to picture as a running notebook of keys and values the model keeps so that it does not reread the whole prompt for every new token, and video fills that notebook faster than any other input. This part does that arithmetic with the sliders doing the counting.
A video is a long sequence
Frames, resolution and rate all multiply
Start with the frame count. A clip — a short stretch of video treated as a single input — that is sampled at the full rate contributes one image's worth of tokens per frame, and the count grows with duration × frame rate × spatial resolution. The reason is that all three factors multiply the sequence: twice the length, twice the frame rate or twice the pixels per side each push the token count up, and they compound rather than add. A three-dimensional tokeniser softens this by compressing time as well as space, where a tokeniser is the rule that chops a continuous signal into discrete units the model can read. Two compressions are stacked here. A causal video VAE — a variational autoencoder, a network that learns to squeeze an input into a small code and then reconstruct it — collapses a small group of frames into one latent frame, a single compressed summary standing in for several raw frames. A spatial compression then shrinks each frame's grid before it is patchified, meaning before it is cut into the square tiles the transformer expects. Both steps help, and the result is still large, because video multiplies the token count along three axes at once where an image multiplies it along two. The bars below put one clip on the same axis as a single image and a page of text, and the gap is what makes video the modality that forces context-window engineering.
The three sliders chain directly, and the chain is worth spelling out because it is the whole cost model in miniature. Frames set the duration, meaning how many raw frames you keep before any compression happens. Resolution sets the spatial grid per frame, so it decides how many patches each kept frame contributes. The frame rate sets how many raw frames the duration contains, because duration is chosen in seconds while the model counts in frames. The token count, the prefill cost — the one-time first pass in which the model reads the whole prompt before generating anything — and the KV footprint all update together, because they are three views of the same sequence length.
One clip against one image and one page of text, drawn with Guide.drawStacked. The clip's token count comes from GenMedia.cost.videoTokens, the image's from Multimodal.cost.imageTokens.
Sampling frames against a budget
Keep every Nth frame, not every frame
The first lever is temporal subsampling, also called frame sampling: instead of feeding the model every frame the camera recorded, you keep one frame out of every few and set the rest aside. Keeping every frame at 24 fps is faithful and unaffordable past a few seconds; keeping every fourth frame is a quarter of the tokens and usually still enough, because most of what changes between adjacent frames is motion the model can infer from the frames it does see. The gap between kept frames is the stride, so a stride of four means one frame kept out of every four. The bars compare three policies for the same clip: every frame, every Nth frame at the stride on the slider, and a fixed token budget that the model is allowed to spend on video — the fixed-size suitcase the clip has to be packed into, however long it turns out to be. The stride does not have to be uniform. Keyframe sampling keeps the frames where the scene changes most and drops the quiet ones, and motion-aware selection does the same by watching how much the picture moves; both are the same idea applied unevenly rather than at a fixed interval. Whatever the schedule, the arithmetic is the same, and it is a direct trade of temporal detail for sequence length.
Token counts for the full-rate clip, the strided clip and the budget, from GenMedia.cost.videoTokens with temporal compression 1 for the full-rate case and the strided frame count for the sampled case.
A stride of 1 is every frame. The budget line is the point of the exercise: choose N so the sampled count fits, and spend what is left on resolution or on keeping a second clip in context.
Temporal position and streaming memory
Order in time, and what to forget
Once a clip is a sequence, the trunk has to know not just which token it is but when it was. Attention is order-blind by construction: it weighs how relevant tokens are to one another but has no built-in sense of where any of them came from, which is why position embeddings exist at all. The useful mental picture is seat numbers written on each token after a shuffle, so the model can tell a patch from the top-left corner from one in the middle. A still-image patch carries a position in space only. A video token carries a position in space and in time, so the position encoding gains a third axis — the temporal axis, the direction running from the start of the clip to its end. The straightforward choice appends a temporal index to the row-and-column indices and encodes all three together. The more careful choice, used by models such as Qwen2-VL, gives each axis its own rotary frequency ranges, where a rotary encoding writes the position in by rotating a token's vector by an angle that depends on where the token sits, so that a token's time offset stays distinguishable from its spatial offset even when both are large. Without an explicit temporal index a patch from the first second and the nine-hundredth look identical to attention, because nothing in the mechanism can tell them apart, which is the video version of the problem position embeddings were invented to fix.
The heatmap shows the sinusoidal temporal encoding as a function of the frame index: each row is a frame, each column a frequency, and the periodic bands are what let attention read "two frames apart" rather than "two tokens apart". "Sinusoidal" here means the pattern is built from sine and cosine waves at several frequencies rather than from a learned table, so the signature of a given frame index is a smooth function of its time, and comparing two signatures is enough to recover how far apart in time the two frames sit.
Qwen3-VL extends the same idea in two directions. It interleaves the M-RoPE frequency bands so the temporal component is spread through the whole embedding rather than living in a reserved slice, which keeps time legible without setting aside a fixed block of dimensions for it and lets the spatial and temporal ranges be traded against each other. It also feeds explicit text-based timestamp tokens — a marker written out as ordinary text, such as <3.2 seconds> — into the sequence, so the model reads elapsed time from a token it can see as well as from the position encoding. The two routes are complementary: position encoding supplies a smooth, continuous sense of how far apart two frames sit, and the timestamp token gives the model a coarse, discrete anchor it can name, so a question that asks what happened after a given moment can be answered by matching text rather than by decoding a position signature.
Sinusoidal temporal position encoding from Multimodal.vit.positional, one row per frame and one column per frequency. The rate of the bands is the model's only handle on elapsed time.
A sliding-window cache over a long stream: each row is a decode step, each column a latent frame, and the bright diagonal band is the window of frames still resident. Built from GenMedia.cost.kvCacheBytes in the readout.
The KV arithmetic
Every token is paid for at every layer
Prefill is a one-off cost, but the key-value cache is not: for every token in context the model stores a key and a value for every layer and every KV head, and it keeps them for as long as the conversation lasts. The KV heads are the parallel attention heads that each hold their own slice of those keys and values. The running-notebook picture is the right one here: every token the model has read stays written down, and the notebook is not torn up between turns. The formula is short and unforgiving. Bytes are $2 \times L \times H_{kv} \times d_{head} \times n_{tokens} \times \text{bytes per value}$, with the leading $2$ for the key and the value, L the number of layers, H_kv the number of KV heads, d_head the width of each head, and n_tokens the context length. At 32 layers, 8 KV heads and 128-dimensional heads in fp16 — fp16 meaning each stored number takes two bytes — that is about 128 KB per token, so a clip of fifty thousand tokens commits roughly six gigabytes of cache — per sequence, and before batching, meaning before the same memory is multiplied across every request the server handles at once. The curve is linear in context length, which is why video is a memory problem as much as a compute one.
KV cache in gigabytes against context length in tokens, from GenMedia.cost.kvCacheBytes via Multimodal.cost.kvBytes. The dashed marker is the current clip's token count.
Prefill FLOPs come from Multimodal.cost.prefillFlops. Attention itself scales with the square of the sequence, so doubling the clip doubles the cache and roughly quadruples the attention work.
Where this shows up
From video QA to generation
Video question answering
A clip and a question enter the same trunk, and the answer is a text token. Everything in this part is on the critical path: sampling decides what the model can see, temporal encoding decides whether it can order events, and the KV cache decides how long a clip the server will accept. Each of the three can be the thing that fails. Miss the moment in sampling and no later stage can recover it; drop the temporal index and "before" and "after" collapse into "somewhere in the clip"; run out of cache and the request is refused rather than answered slowly. The spatiotemporal encoder — the network that reads space and time together rather than treating each frame as an independent picture — is the video case of the patch transformer.
Video as token prediction
The reverse direction, video out, is the same latent token stream decoded instead of encoded: a video tokeniser supplies the vocabulary, meaning the fixed set of discrete units the model is allowed to emit, and a transformer predicts it. The word "latent" is doing real work here — these are compressed codes standing in for patches of pixels, not pixels themselves, so the model predicts in a compact space and a decoder turns the result back into frames. That is the generative-media volume's video chapter, and the sampling and cache arithmetic are shared between the two directions: whichever way the tokens flow, the sequence length and the memory it implies are the same.
The next part takes the same questions to the modality with no frames at all. Audio arrives as a continuous waveform — a pressure signal that varies smoothly in time, with no natural grid the way video has frames — whose token rate and temporal structure are set by the codec rather than by a frame rate, where the codec is the compressor that decides how that waveform is chopped into discrete units. That difference matters because there is no obvious stride to turn down when the budget is tight: the compression decision has already been made, by the codec, before the trunk sees anything. The choice between an encoder's transcript and native audio tokens — the codec's own units, kept as sound rather than converted into written words — changes what the trunk can hear, because a transcript has already thrown away tone, emphasis and everything else a listener would notice but a reader would not.
Further reading
These papers mark the path from treating video as a stack of images to treating it as one long sequence with a temporal axis, and they supply the position and caching machinery the arithmetic above depends on. Read in order, they move from factorising space and time in a small video transformer, through attention designs that divide the work between the two axes, to a token-based generator, and finally to the position encoding and long-context framing that today's models rely on. The thread running through them is the same question the sliders above answer: how do you preserve the order of a long sequence without paying separately for every frame?
- Anurag Arnab, Mostafa Dehghani, Georg Heigold and coauthors, "ViViT: A Video Vision Transformer", 2021 — factorising space and time in a video transformer.
- Gedas Bertasius, Heng Wang and Lorenzo Torresani, "Is Space-Time Attention All You Need for Video Understanding?", 2021 — TimeSformer and the divided-attention variants.
- Dan Kondratyuk, Lijun Yu, Xiuye Gu and coauthors, "VideoPoet: A Large Language Model for Zero-Shot Video Generation", 2023 — video generated as a token sequence by a decoder-only model.
- Peng Wang, Shuai Bai, Sinan Tan and coauthors, "Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution", 2024 — multidimensional rotary position encoding for image and video tokens.
- Shuai Bai, Yuxuan Cai, Ruizhe Chen and coauthors, "Qwen3-VL Technical Report", 2025 — interleaved M-RoPE and text-based timestamps for temporal structure.
- OpenAI, "Video generation models as world simulators", 2024 — the patched-spacetime and variable-duration framing that makes clips sequences.
Cheat sheet
| Term | Meaning here |
|---|---|
| Frames × tokens | A clip's length is duration × frame rate × tokens per frame |
| 3-D tokeniser | Spatial and temporal compression before patchification; a VAE collapses frame groups |
| Sampling stride | Keep every Nth frame; the direct lever on token count |
| Token budget | A fixed number of video tokens the model may spend, whatever the clip |
| Temporal position | A third position axis, or a third rotary frequency range, so order survives |
| Streaming memory | Sliding windows, summaries and attention sinks: a policy for what to forget |
| KV per token | $2 \times L \times H_{kv} \times d_{head} \times \text{bytes}$ — about 128 KB/token at 32×8×128 in fp16 |
| Prefill scaling | Attention work grows with the square of the sequence; the cache grows linearly |