Resolution and the token budget
An image is not one token. A token is the smallest unit a language model reads and writes — a small chunk of the signal mapped to an integer id — and a picture has to be broken into many of them before the model can look at it at all. It is as many tokens as the patch grid says it is, where the patch grid is the chequerboard of equal squares the image is cut into before anything else happens. Every one of those tokens behaves like a word in the prompt. Each one lengthens the attention matrix, the table recording how much every token looks at every other token, whose cost grows with the square of the sequence length. Each one occupies a slot in the KV cache, the running notebook of keys and values the model keeps so that it does not reread the prompt on every new word. And each one is paid for at prefill, the single first pass in which the model reads the whole prompt before it generates anything, and then again on every decode step that follows. Resolution, here, means how many pixels across the image is sampled: a finer sampling gives smaller pieces and more of them, so more tokens. Because the model can hold only a fixed number of tokens at once, the image's share of that allowance is what this part calls the token budget, and it is the fixed-size suitcase every high-resolution request has to be packed into. A vision-language model is therefore always trading spatial detail against sequence length, and the three schemes in use — a fixed $224^2$ resize, AnyRes tiling, and native dynamic resolution — are three answers to that trade. The first shrinks every image to one small square. The second cuts a large image into encoder-sized tiles and adds one small view of the whole. The third skips the square and patches the image at its own size, letting the token count float. This part follows the arithmetic from image size to tile count to token count to FLOPs and bytes.
Why resolution costs
Attention is quadratic in the sequence, and an image is a long sequence
Start with the arithmetic. A patch transformer is the vision encoder that cuts an image into a grid of square patches and treats each patch as a token, so before anything else the image is chopped into equal squares. It turns an $H \times H$ image into $(H/P)^2$ tokens at patch size $P$, because the grid is $H/P$ patches across and $H/P$ patches down, and each square becomes one token. Dicing an onion works the same way: a finer dice gives more, smaller pieces, and it is the number of pieces you pay for rather than the mass of the onion. At $224^2$ with $P = 14$ that is $16 \times 16 = 256$ tokens; at $448^2$ it is $1024$; at $896^2$ it is $4096$. Doubling the side of the image quadruples the token count, and that is the quadratic the rest of the part keeps running into. The self-attention inside the language model costs time and memory proportional to the square of the total sequence length, so the image's contribution grows quadratically while the text's stays small. Why does that matter so much? Because the context window — the fixed maximum number of tokens the model can hold in front of it at once — is a hard ceiling, and an image can spend a large part of it. An image roughly the size of a page of text at $448^2$ is already several paragraphs of prompt, and at native phone-camera resolution it is the whole context window.
The reason to pay anyway is that low resolution destroys exactly the information multimodal models are most often asked about. A $224^2$ resize of a dense document is unreadable; a chart's axis labels vanish; small objects become single pixels. Text is the sharpest example of the loss. A character is only a few pixels wide at best, so downscaling a page into a small square averages the strokes of each letter into grey mush, and no amount of language ability can recover a word that is no longer present in the pixels. The whole design space of multimodal resolution is an attempt to keep the fine detail that high resolution provides while refusing to pay the full sequence cost of carrying every patch separately. The knobs are the patch size, the tiling policy, and a learned or fixed merge that groups neighbouring patches into one token. Tiling means splitting the image into several fixed-size pieces and encoding each piece on its own; a merge means combining a small block of neighbouring patches into a single token after the encoder has produced them, which is the compression lever the last step of this part prices. All three knobs trade the same two quantities against each other, and every scheme in the literature is a different setting of them.
The rest of the part builds that arithmetic one step at a time, and then prices it. The final step turns the token count into prefill FLOPs and KV cache bytes, which is where LLM Serving, Part 18 picks the story up: an image inverts the usual prefill-heavy, decode-light serving profile, because it adds a large amount of prefill work that then persists as cache on every subsequent decode step. Two units appear there and are worth naming now. FLOPs, short for floating-point operations, are the standard way to count how much arithmetic a model performs, so the prefill pass's FLOP count is the work of reading every prompt token once. Bytes are the standard way to measure the memory the KV cache occupies, and that cache is not rebuilt for each new word — it is reread, which is why a picture is not a one-off cost but a permanent tax on the throughput of everything generated after it.
Fixed patches
Resize everything to one square and be done
The simplest policy is the one CLIP was trained under: resize every image to a fixed square — typically $224 \times 224$ — and patchify that. Patchify means cut the square into the grid of patches the encoder expects and flatten each patch into a vector; because the square is always the same size, so is the grid. The token count is then a constant, independent of the input, which makes batching trivial and the language model's context predictable. Batching is running many requests through the model at once, and a constant token count means every image occupies exactly the same number of slots, so requests can be grouped without padding arithmetic. The price is paid in the resize. Aspect ratio, the ratio of an image's width to its height, is discarded, so a wide panorama is squashed, a tall page is shrunk until its text is a blur, and the model sees a downsampled version of the scene that may no longer contain the thing it is being asked about. The flat curve below is that policy's token count; the rising curve is what the same encoder would produce if it were allowed to keep the input's native size.
The gap between the two curves is the entire motivation for the next two steps. There are only two ways to behave differently. One way to close it is to stop resizing and let the token count grow, which is native dynamic resolution; the other is to keep the input large but spend fewer tokens on it, which is tiling and merging. Notice that the two answers trade in different currencies. Native resolution spends context to buy detail, while tiling and merging spend extra work on the vision side to buy detail without spending as much context. Neither is free, and the rest of the part is about seeing exactly what each one costs.
Image tokens against input side for the fixed-$224^2$ resize (flat) and for native patchification at the same patch size (rising). The dashed line is the classic $224$ input.
Fixed $224^2$ tokens come from Multimodal.cost.imageTokens; the rising curve is the native count at the same patch size.
AnyRes tiling
Cut the image into encoder-sized tiles, plus a thumbnail
AnyRes, the scheme LLaVA-NeXT made standard, refuses to choose between detail and a fixed encoder input. Instead of shrinking the whole image to fit the encoder, it splits a high-resolution image into a grid of tiles that are each the size the encoder was trained on — $336 \times 336$, say — encodes every tile separately at full resolution, and additionally encodes a downsampled version of the whole image as a thumbnail. Each tile is therefore seen in the detail it deserves, because nothing inside it was downscaled; the thumbnail is the small global view that shows how the tiles fit together, and it is the only place the whole page appears at once. The language model receives one block of vision tokens per tile plus the thumbnail's block, so the token count is $(\text{tiles} + 1)$ times the per-tile count. A $672 \times 336$ image with $336$ tiles is exactly two tiles plus the thumbnail; a $1344 \times 672$ image is eight plus one. The point of the grid is that the number of tiles is decided by the image's area, so a page of dense text can be given as many tiles as it needs while each individual tile stays at the resolution the encoder understands.
The structure has two clear properties, and each is a cost to keep in mind. First, the token count scales with the number of tiles, which scales with area, so a very large image still costs a lot of sequence. A picture twice as wide and twice as tall needs four times the tiles and therefore roughly four times the tokens, exactly as native resolution would. Second, the tiling is a hard cut: an object straddling a tile boundary is seen in two halves, and the thumbnail is the only view that puts it back together. Attention inside one tile cannot see the neighbouring tile, so a face split by a seam is never present whole in any single block of tokens; the low-resolution thumbnail is what tells the model that the two halves belong to one object. The variation of AnyRes used in practice (the "grid pinpoints") picks the tiling layout whose aspect ratio most closely matches the image, which keeps the token count from growing on one axis more than necessary. Choosing the layout well matters, because a mismatched grid wastes tiles on empty margins, and every wasted tile is a full block of tokens that adds nothing.
The tiling layout from Multimodal.tiles.tiles, with the global thumbnail drawn to scale on the right. Each tile contributes its own block of vision tokens.
Native dynamic resolution
Keep the aspect ratio, let the token count float
Native dynamic resolution, the scheme behind Qwen-VL and InternVL, skips the resizing and the tiling grid entirely. The image is patched at its own size and aspect ratio, so a tall, narrow page is handled as a tall, narrow page rather than being forced into a square, and the token count is area divided by patch area, adjusted by the merge factor: a bigger image produces more tokens, and the merge can pull that count back down. The encoder is modified to accept variable-length patch sequences, which means it can no longer assume a fixed grid of, say, $16 \times 16$ tokens. To keep the model oriented, these encoders usually use a two-dimensional rotary position embedding, a way of writing position into each token by rotating its vector by an angle that depends on the token's row and column. Position embeddings are the addresses that tell an otherwise order-blind transformer where each patch came from; a two-dimensional version encodes both coordinates at once, and because the address is a continuous function of row and column rather than a fixed lookup table, the model can generalise to grids it was never trained on. There is no thumbnail and no boundary cut, so nothing is lost at the seams.
What replaces the fixed budget is a decision per image: the pipeline picks a resolution, sometimes rounded up to the nearest multiple of the patch and merge sizes, based on the task and the context available. The rounding matters because the grid has to divide evenly into patches and then into merge blocks; a resolution that leaves a partial row would otherwise have to be dropped or padded. Document and OCR requests get many tokens; casual descriptions get few. OCR stands for optical character recognition, the task of reading text out of an image, and it is the most resolution-hungry request there is, because a single character may occupy only a handful of pixels. The trade is now made at request time rather than fixed once in training, which is the design's main advantage: the same model can spend a couple of hundred tokens on a casual photo and several thousand on a photographed contract. The plot below shows the native count rising with the square of the side, and the effect of the merge from the next step as a dashed line that grows much more slowly.
A second lever sits inside the encoder rather than at the token count. Qwen2.5-VL scales its vision transformer with windowed self-attention: most ViT layers attend only inside local $112 \times 112$ windows — a fixed neighbourhood of patches, so each patch sees its neighbours directly and nothing farther away — and full attention, the expensive kind that compares every patch with every other, runs at just a few layers. Because the quadratic term in attention grows with the square of the whole token grid, keeping full attention to a handful of layers means the encoder never pays that price over the whole grid at once, so the cost of a large image stops exploding as it scales up. This is a different answer from tiling, which cuts the image into pieces and accepts the seams, and from merging, which shrinks the sequence after the encoder has already paid for it; windowed attention leaves the token grid intact and changes what each layer is allowed to read. Where this shows up is Pixtral 12B, which handles several images in one sequence at their native resolution: each image keeps its own aspect ratio and its own token count rather than being forced into a common square, so a wide panorama, a tall page and a small photo can sit in the same prompt side by side, and the model sees several images of different shapes in one sequence rather than one image at a time.
Native token count against image side, with and without the merge, from Multimodal.tiles.nativeTokens. The dashed line marks the current input side.
Doubling the side quadruples the token count. That is the quadratic that the merge exists to soften.
Compressing the tokens
Four patches, one token
A merge layer groups a small block of neighbouring patch tokens into a single token before they reach the language model, which cuts the sequence by the square of the block size. This is the dicing-an-onion move in the other direction: rather than handing every grain of the dice to the language model, you pack a $2 \times 2$ cluster of grains into one larger parcel. A $2 \times 2$ merge divides the token count by four; a $4 \times 4$ merge by sixteen. A real model ships at exactly that number: InternVL3.5's dynamic-resolution path encodes each image patch as 1024 visual tokens and pixel-shuffles them to 256 before the language model sees them, the fourfold reduction the $2 \times 2$ block describes. Pooling does this by averaging the block, which is cheap and destroys fine detail. Averaging is a smoothing operation: if one patch carries a small, unique mark and its three neighbours carry background, the mean keeps the background and washes the mark out, and the information is gone before the language model ever sees it. Pixel-shuffle does it by concatenating the block's channels instead of averaging them: a $2 \times 2$ block of $C$-dimensional patches becomes one token of dimension $4C$, so no information is thrown away at the merge itself — the four patches' features are laid side by side into a wider vector, and the language model is asked to work with wider vectors instead of more of them. Perceiver-style resampling, which is the Q-Former's trick, is the learned version of the same idea, replacing the fixed block with learned queries that decide what to keep. A fixed merge treats every region of the image as equally important; a learned one can spend its tokens where the content actually is.
The bars below compare the options at the current input size. The token reduction is real and large, and it is the single most effective lever a practitioner has for fitting high-resolution input into a fixed context. A $4 \times 4$ merge cuts the sequence to a sixteenth of its length, which is the difference between a document that fits in the context window and one that does not. The caveat is the same one that applies to any compression: the detail that disappears at the merge cannot be recovered downstream. Once four patches have been averaged into one vector, no later layer can tell what the individual patches held, so the factor has to be chosen with the task in mind rather than by picking the largest one that fits.
Image tokens at the current input side for each pooling option, from Multimodal.cost.imageTokens. Pixel-shuffle matches pooling on count but keeps the block's channels.
A learned resampler sits at the same point in the pipeline; it just decides the compression from the image rather than by a fixed block.
Where this shows up
Tokens become FLOPs and bytes
The token count is the intermediate quantity; what a serving system feels is the prefill cost and the KV cache. Prefill runs the full attention and MLP stack over every prompt token at once, so its FLOP count grows with the sequence linearly in the projections and quadratically in attention. The MLP is the feed-forward block that transforms each token on its own, and the projections are the learned matrices that map a token into its query, key and value; both are per-token work, so they cost a fixed amount per token, whereas attention compares every token with every other and therefore grows with the square of the sequence. The KV cache then stores a key and a value for every token at every layer, so an image's tokens stay resident and are re-read on every decode step. That is the asymmetry to hold on to: prefill is a one-time arithmetic cost, while the cache is a permanent memory cost that is touched again for each new word the model writes. Both are computed here with Multimodal.cost.prefillFlops and Multimodal.cost.kvBytes.
The plots make the inversion that LLM Serving, Part 18 works out visible: a text-only request is often decode-dominated, but a request with a high-resolution image is a large prefill followed by cache-heavy decode, which is the opposite shape. A text-only prompt is short, so the model spends most of its time generating words one at a time, each step reading a small cache; a high-resolution image makes the prompt long, so the one big read at the start dominates, and every later step has to drag the whole image-shaped cache along with it. That is why multimodal serving tends to be memory-bandwidth bound — the bottleneck is how fast keys and values can be moved from memory, not how many FLOPs the arithmetic needs — and why compression at the vision side is worth more than it looks.
Prefill FLOPs and KV cache bytes against total prompt tokens, with the current image's contribution marked. The text prompt is assumed to be 64 tokens.
Head dimension is fixed at 128. Reducing KV heads with grouped-query attention cuts the cache in proportion without touching the FLOPs.
Further reading
These papers trace the path from a fixed $224^2$ input to native dynamic resolution, and the serving analysis on the other side of the token budget. Read them in order and the design pressure becomes visible: each one is reacting to the resolution limit of the paper before it, and the serving reference is there to show what the accumulated token count costs once the model is actually deployed.
- Haotian Liu, Chunyuan Li, Yuheng Li and Yong Jae Lee, "Improved Baselines with Visual Instruction Tuning", 2023 — LLaVA-1.5, and the AnyRes tiling scheme that followed.
- Haotian Liu, Chunyuan Li, Bo Li and Yong Jae Lee, "LLaVA-NeXT: Improved reasoning, OCR, and world knowledge", 2024 — grid pinpointing and the high-resolution recipe for documents and charts.
- Shuai Bai, Keqin Chen, Xuejing Liu and coauthors, "Qwen2.5-VL Technical Report", 2025 — windowed ViT attention and native dynamic resolution.
- Weiyun Wang, Zhangwei Gao, Lixin Gu and coauthors, "InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency", 2025 — 1024-to-256 pixel shuffle and the dynamic high-resolution path.
- Pravesh Agrawal, Szymon Antoniak, Emma Bou Hanna and coauthors, "Pixtral 12B", 2024 — native resolution with multiple images per sequence.
- Zhiliang Peng, Wenhui Wang, Li Dong and coauthors, "Kosmos-2: Grounding Multimodal Large Language Models to the World", 2023 — spatial tokens and the resolution questions that grounding raises.
- LLM Serving, Part 18: Workloads — the prefill-and-decode accounting that this part's final demo feeds into.
Cheat sheet
| Term | Meaning here |
|---|---|
| Fixed patches | Resize to one square, patchify, constant token count; detail lost in the resize |
| Token count | $(H/P)^2$ before merging; doubles the side, quadruples the tokens |
| AnyRes | Encode a grid of tiles plus a whole-image thumbnail; tokens scale with tile count |
| Grid pinpoints | Choosing the tile layout whose aspect ratio best matches the image |
| Native resolution | Patch at the image's own size and aspect ratio, with 2-D position embeddings |
| 2-D RoPE | Rotary position encoding with separate row and column phases, so unseen grids are usable |
| Merge / pooling | Group a $k \times k$ block of patches into one token, dividing the budget by $k^2$ |
| Pixel-shuffle | Repack a block into channels instead of averaging, keeping detail at the same token count |
| Pruning | Score patches and drop the uninformative ones; irregular but detail-preserving |
| Prefill vs decode | An image adds a large one-time prefill and a permanent KV cache |