Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

A video is a long sequence

Frames, resolution and rate all multiply

Start with the frame count. A clip — a short stretch of video treated as a single input — that is sampled at the full rate contributes one image's worth of tokens per frame, and the count grows with duration × frame rate × spatial resolution. The reason is that all three factors multiply the sequence: twice the length, twice the frame rate or twice the pixels per side each push the token count up, and they compound rather than add. A three-dimensional tokeniser softens this by compressing time as well as space, where a tokeniser is the rule that chops a continuous signal into discrete units the model can read. Two compressions are stacked here. A causal video VAE — a variational autoencoder, a network that learns to squeeze an input into a small code and then reconstruct it — collapses a small group of frames into one latent frame, a single compressed summary standing in for several raw frames. A spatial compression then shrinks each frame's grid before it is patchified, meaning before it is cut into the square tiles the transformer expects. Both steps help, and the result is still large, because video multiplies the token count along three axes at once where an image multiplies it along two. The bars below put one clip on the same axis as a single image and a page of text, and the gap is what makes video the modality that forces context-window engineering.

The three sliders chain directly, and the chain is worth spelling out because it is the whole cost model in miniature. Frames set the duration, meaning how many raw frames you keep before any compression happens. Resolution sets the spatial grid per frame, so it decides how many patches each kept frame contributes. The frame rate sets how many raw frames the duration contains, because duration is chosen in seconds while the model counts in frames. The token count, the prefill cost — the one-time first pass in which the model reads the whole prompt before generating anything — and the KV footprint all update together, because they are three views of the same sequence length.

One clip against one image and one page of text, drawn with Guide.drawStacked. The clip's token count comes from GenMedia.cost.videoTokens, the image's from Multimodal.cost.imageTokens.

💡 By the end of this part you'll be able to turn frames, resolution and rate into a token count, decide a sampling stride from a token budget, explain temporal position encoding and streaming memory, and say how much KV cache a clip commits the server to. Those four skills are one chain: sampling sets the token count, the budget caps it, position encoding keeps the order straight, and the cache is the memory bill for every token that stays in context.
2

Sampling frames against a budget

Keep every Nth frame, not every frame

The first lever is temporal subsampling, also called frame sampling: instead of feeding the model every frame the camera recorded, you keep one frame out of every few and set the rest aside. Keeping every frame at 24 fps is faithful and unaffordable past a few seconds; keeping every fourth frame is a quarter of the tokens and usually still enough, because most of what changes between adjacent frames is motion the model can infer from the frames it does see. The gap between kept frames is the stride, so a stride of four means one frame kept out of every four. The bars compare three policies for the same clip: every frame, every Nth frame at the stride on the slider, and a fixed token budget that the model is allowed to spend on video — the fixed-size suitcase the clip has to be packed into, however long it turns out to be. The stride does not have to be uniform. Keyframe sampling keeps the frames where the scene changes most and drops the quiet ones, and motion-aware selection does the same by watching how much the picture moves; both are the same idea applied unevenly rather than at a fixed interval. Whatever the schedule, the arithmetic is the same, and it is a direct trade of temporal detail for sequence length.

Token counts for the full-rate clip, the strided clip and the budget, from GenMedia.cost.videoTokens with temporal compression 1 for the full-rate case and the strided frame count for the sampled case.

A stride of 1 is every frame. The budget line is the point of the exercise: choose N so the sampled count fits, and spend what is left on resolution or on keeping a second clip in context.

⚠ Uniform subsampling loses exactly the wrong frames. A fast gesture or a single-word subtitle can live entirely between two sampled frames, and the model is not wrong about what it saw — it never saw the moment at all. Production video models therefore mix a low frame rate for the whole clip with a denser sample around detected motion or around regions the question mentions — the same token budget spent unevenly.
3

Temporal position and streaming memory

Order in time, and what to forget

Once a clip is a sequence, the trunk has to know not just which token it is but when it was. Attention is order-blind by construction: it weighs how relevant tokens are to one another but has no built-in sense of where any of them came from, which is why position embeddings exist at all. The useful mental picture is seat numbers written on each token after a shuffle, so the model can tell a patch from the top-left corner from one in the middle. A still-image patch carries a position in space only. A video token carries a position in space and in time, so the position encoding gains a third axis — the temporal axis, the direction running from the start of the clip to its end. The straightforward choice appends a temporal index to the row-and-column indices and encodes all three together. The more careful choice, used by models such as Qwen2-VL, gives each axis its own rotary frequency ranges, where a rotary encoding writes the position in by rotating a token's vector by an angle that depends on where the token sits, so that a token's time offset stays distinguishable from its spatial offset even when both are large. Without an explicit temporal index a patch from the first second and the nine-hundredth look identical to attention, because nothing in the mechanism can tell them apart, which is the video version of the problem position embeddings were invented to fix.

The heatmap shows the sinusoidal temporal encoding as a function of the frame index: each row is a frame, each column a frequency, and the periodic bands are what let attention read "two frames apart" rather than "two tokens apart". "Sinusoidal" here means the pattern is built from sine and cosine waves at several frequencies rather than from a learned table, so the signature of a given frame index is a smooth function of its time, and comparing two signatures is enough to recover how far apart in time the two frames sit.

Qwen3-VL extends the same idea in two directions. It interleaves the M-RoPE frequency bands so the temporal component is spread through the whole embedding rather than living in a reserved slice, which keeps time legible without setting aside a fixed block of dimensions for it and lets the spatial and temporal ranges be traded against each other. It also feeds explicit text-based timestamp tokens — a marker written out as ordinary text, such as <3.2 seconds> — into the sequence, so the model reads elapsed time from a token it can see as well as from the position encoding. The two routes are complementary: position encoding supplies a smooth, continuous sense of how far apart two frames sit, and the timestamp token gives the model a coarse, discrete anchor it can name, so a question that asks what happened after a given moment can be answered by matching text rather than by decoding a position signature.

Sinusoidal temporal position encoding from Multimodal.vit.positional, one row per frame and one column per frequency. The rate of the bands is the model's only handle on elapsed time.

A sliding-window cache over a long stream: each row is a decode step, each column a latent frame, and the bright diagonal band is the window of frames still resident. Built from GenMedia.cost.kvCacheBytes in the readout.

💡 A stream has no natural end, so a model that keeps every token it has ever seen will eventually run out of memory. Streaming memory is the discipline of holding state over an input that keeps arriving, and it is easier to picture as passing notes in a relay: each stage hands a summary forward instead of every stage carrying the whole history. The standard answers are a sliding window that keeps the most recent frames, a summary that compresses older ones into fewer tokens, or attention sinks that retain a few early tokens as anchors, where an attention sink is a token kept on purpose because later attention tends to lean on it. All three are policy choices about what a video model is allowed to forget.
4

The KV arithmetic

Every token is paid for at every layer

Prefill is a one-off cost, but the key-value cache is not: for every token in context the model stores a key and a value for every layer and every KV head, and it keeps them for as long as the conversation lasts. The KV heads are the parallel attention heads that each hold their own slice of those keys and values. The running-notebook picture is the right one here: every token the model has read stays written down, and the notebook is not torn up between turns. The formula is short and unforgiving. Bytes are $2 \times L \times H_{kv} \times d_{head} \times n_{tokens} \times \text{bytes per value}$, with the leading $2$ for the key and the value, L the number of layers, H_kv the number of KV heads, d_head the width of each head, and n_tokens the context length. At 32 layers, 8 KV heads and 128-dimensional heads in fp16 — fp16 meaning each stored number takes two bytes — that is about 128 KB per token, so a clip of fifty thousand tokens commits roughly six gigabytes of cache — per sequence, and before batching, meaning before the same memory is multiplied across every request the server handles at once. The curve is linear in context length, which is why video is a memory problem as much as a compute one.

KV cache in gigabytes against context length in tokens, from GenMedia.cost.kvCacheBytes via Multimodal.cost.kvBytes. The dashed marker is the current clip's token count.

Prefill FLOPs come from Multimodal.cost.prefillFlops. Attention itself scales with the square of the sequence, so doubling the clip doubles the cache and roughly quadruples the attention work.

⚠ The cache is the reason video and long context are the same conversation. Long context here just means a context window stretched to tens or hundreds of thousands of tokens, and the cache does not care what filled it. A serving stack that can hold a hundred thousand tokens of text can hold the same count of video tokens; nothing about the modality changes the arithmetic. What video changes is how quickly you reach that number, which is why the sampling stride from step 2 is really a memory decision in disguise.
5

Where this shows up

From video QA to generation

Understanding

Video question answering

A clip and a question enter the same trunk, and the answer is a text token. Everything in this part is on the critical path: sampling decides what the model can see, temporal encoding decides whether it can order events, and the KV cache decides how long a clip the server will accept. Each of the three can be the thing that fails. Miss the moment in sampling and no later stage can recover it; drop the temporal index and "before" and "after" collapse into "somewhere in the clip"; run out of cache and the request is refused rather than answered slowly. The spatiotemporal encoder — the network that reads space and time together rather than treating each frame as an independent picture — is the video case of the patch transformer.

Generation

Video as token prediction

The reverse direction, video out, is the same latent token stream decoded instead of encoded: a video tokeniser supplies the vocabulary, meaning the fixed set of discrete units the model is allowed to emit, and a transformer predicts it. The word "latent" is doing real work here — these are compressed codes standing in for patches of pixels, not pixels themselves, so the model predicts in a compact space and a decoder turns the result back into frames. That is the generative-media volume's video chapter, and the sampling and cache arithmetic are shared between the two directions: whichever way the tokens flow, the sequence length and the memory it implies are the same.

The next part takes the same questions to the modality with no frames at all. Audio arrives as a continuous waveform — a pressure signal that varies smoothly in time, with no natural grid the way video has frames — whose token rate and temporal structure are set by the codec rather than by a frame rate, where the codec is the compressor that decides how that waveform is chopped into discrete units. That difference matters because there is no obvious stride to turn down when the budget is tight: the compression decision has already been made, by the codec, before the trunk sees anything. The choice between an encoder's transcript and native audio tokens — the codec's own units, kept as sound rather than converted into written words — changes what the trunk can hear, because a transcript has already thrown away tone, emphasis and everything else a listener would notice but a reader would not.

Further reading

These papers mark the path from treating video as a stack of images to treating it as one long sequence with a temporal axis, and they supply the position and caching machinery the arithmetic above depends on. Read in order, they move from factorising space and time in a small video transformer, through attention designs that divide the work between the two axes, to a token-based generator, and finally to the position encoding and long-context framing that today's models rely on. The thread running through them is the same question the sliders above answer: how do you preserve the order of a long sequence without paying separately for every frame?

Cheat sheet

TermMeaning here
Frames × tokensA clip's length is duration × frame rate × tokens per frame
3-D tokeniserSpatial and temporal compression before patchification; a VAE collapses frame groups
Sampling strideKeep every Nth frame; the direct lever on token count
Token budgetA fixed number of video tokens the model may spend, whatever the clip
Temporal positionA third position axis, or a third rotary frequency range, so order survives
Streaming memorySliding windows, summaries and attention sinks: a policy for what to forget
KV per token$2 \times L \times H_{kv} \times d_{head} \times \text{bytes}$ — about 128 KB/token at 32×8×128 in fp16
Prefill scalingAttention work grows with the square of the sequence; the cache grows linearly
7

Check your understanding

0/4 answered