Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

Why resolution costs

Attention is quadratic in the sequence, and an image is a long sequence

Start with the arithmetic. A patch transformer is the vision encoder that cuts an image into a grid of square patches and treats each patch as a token, so before anything else the image is chopped into equal squares. It turns an $H \times H$ image into $(H/P)^2$ tokens at patch size $P$, because the grid is $H/P$ patches across and $H/P$ patches down, and each square becomes one token. Dicing an onion works the same way: a finer dice gives more, smaller pieces, and it is the number of pieces you pay for rather than the mass of the onion. At $224^2$ with $P = 14$ that is $16 \times 16 = 256$ tokens; at $448^2$ it is $1024$; at $896^2$ it is $4096$. Doubling the side of the image quadruples the token count, and that is the quadratic the rest of the part keeps running into. The self-attention inside the language model costs time and memory proportional to the square of the total sequence length, so the image's contribution grows quadratically while the text's stays small. Why does that matter so much? Because the context window — the fixed maximum number of tokens the model can hold in front of it at once — is a hard ceiling, and an image can spend a large part of it. An image roughly the size of a page of text at $448^2$ is already several paragraphs of prompt, and at native phone-camera resolution it is the whole context window.

The reason to pay anyway is that low resolution destroys exactly the information multimodal models are most often asked about. A $224^2$ resize of a dense document is unreadable; a chart's axis labels vanish; small objects become single pixels. Text is the sharpest example of the loss. A character is only a few pixels wide at best, so downscaling a page into a small square averages the strokes of each letter into grey mush, and no amount of language ability can recover a word that is no longer present in the pixels. The whole design space of multimodal resolution is an attempt to keep the fine detail that high resolution provides while refusing to pay the full sequence cost of carrying every patch separately. The knobs are the patch size, the tiling policy, and a learned or fixed merge that groups neighbouring patches into one token. Tiling means splitting the image into several fixed-size pieces and encoding each piece on its own; a merge means combining a small block of neighbouring patches into a single token after the encoder has produced them, which is the compression lever the last step of this part prices. All three knobs trade the same two quantities against each other, and every scheme in the literature is a different setting of them.

The rest of the part builds that arithmetic one step at a time, and then prices it. The final step turns the token count into prefill FLOPs and KV cache bytes, which is where LLM Serving, Part 18 picks the story up: an image inverts the usual prefill-heavy, decode-light serving profile, because it adds a large amount of prefill work that then persists as cache on every subsequent decode step. Two units appear there and are worth naming now. FLOPs, short for floating-point operations, are the standard way to count how much arithmetic a model performs, so the prefill pass's FLOP count is the work of reading every prompt token once. Bytes are the standard way to measure the memory the KV cache occupies, and that cache is not rebuilt for each new word — it is reread, which is why a picture is not a one-off cost but a permanent tax on the throughput of everything generated after it.

💡 By the end of this part you'll be able to compute the token count for a fixed-patch, AnyRes or native-resolution input, say what a 2×2 merge does to the budget, explain why OCR and document tasks force high resolution, and translate an image's tokens into prefill FLOPs and KV bytes. Each of those is one link in the same chain from pixels to tokens to compute.
2

Fixed patches

Resize everything to one square and be done

The simplest policy is the one CLIP was trained under: resize every image to a fixed square — typically $224 \times 224$ — and patchify that. Patchify means cut the square into the grid of patches the encoder expects and flatten each patch into a vector; because the square is always the same size, so is the grid. The token count is then a constant, independent of the input, which makes batching trivial and the language model's context predictable. Batching is running many requests through the model at once, and a constant token count means every image occupies exactly the same number of slots, so requests can be grouped without padding arithmetic. The price is paid in the resize. Aspect ratio, the ratio of an image's width to its height, is discarded, so a wide panorama is squashed, a tall page is shrunk until its text is a blur, and the model sees a downsampled version of the scene that may no longer contain the thing it is being asked about. The flat curve below is that policy's token count; the rising curve is what the same encoder would produce if it were allowed to keep the input's native size.

The gap between the two curves is the entire motivation for the next two steps. There are only two ways to behave differently. One way to close it is to stop resizing and let the token count grow, which is native dynamic resolution; the other is to keep the input large but spend fewer tokens on it, which is tiling and merging. Notice that the two answers trade in different currencies. Native resolution spends context to buy detail, while tiling and merging spend extra work on the vision side to buy detail without spending as much context. Neither is free, and the rest of the part is about seeing exactly what each one costs.

Image tokens against input side for the fixed-$224^2$ resize (flat) and for native patchification at the same patch size (rising). The dashed line is the classic $224$ input.

Fixed $224^2$ tokens come from Multimodal.cost.imageTokens; the rising curve is the native count at the same patch size.

⚠ The resize is the model's blind spot. If a task fails at low resolution it can be a data problem, a prompt problem, or because the pixels carrying the answer did not survive the downscale. Before blaming the language model, check what the vision tower was actually given: if the answer is not in the pixels the encoder received, no prompt or model change can put it back.
3

AnyRes tiling

Cut the image into encoder-sized tiles, plus a thumbnail

AnyRes, the scheme LLaVA-NeXT made standard, refuses to choose between detail and a fixed encoder input. Instead of shrinking the whole image to fit the encoder, it splits a high-resolution image into a grid of tiles that are each the size the encoder was trained on — $336 \times 336$, say — encodes every tile separately at full resolution, and additionally encodes a downsampled version of the whole image as a thumbnail. Each tile is therefore seen in the detail it deserves, because nothing inside it was downscaled; the thumbnail is the small global view that shows how the tiles fit together, and it is the only place the whole page appears at once. The language model receives one block of vision tokens per tile plus the thumbnail's block, so the token count is $(\text{tiles} + 1)$ times the per-tile count. A $672 \times 336$ image with $336$ tiles is exactly two tiles plus the thumbnail; a $1344 \times 672$ image is eight plus one. The point of the grid is that the number of tiles is decided by the image's area, so a page of dense text can be given as many tiles as it needs while each individual tile stays at the resolution the encoder understands.

The structure has two clear properties, and each is a cost to keep in mind. First, the token count scales with the number of tiles, which scales with area, so a very large image still costs a lot of sequence. A picture twice as wide and twice as tall needs four times the tiles and therefore roughly four times the tokens, exactly as native resolution would. Second, the tiling is a hard cut: an object straddling a tile boundary is seen in two halves, and the thumbnail is the only view that puts it back together. Attention inside one tile cannot see the neighbouring tile, so a face split by a seam is never present whole in any single block of tokens; the low-resolution thumbnail is what tells the model that the two halves belong to one object. The variation of AnyRes used in practice (the "grid pinpoints") picks the tiling layout whose aspect ratio most closely matches the image, which keeps the token count from growing on one axis more than necessary. Choosing the layout well matters, because a mismatched grid wastes tiles on empty margins, and every wasted tile is a full block of tokens that adds nothing.

The tiling layout from Multimodal.tiles.tiles, with the global thumbnail drawn to scale on the right. Each tile contributes its own block of vision tokens.

💡 The thumbnail is what makes the grid global. Per-tile tokens are local by construction, because attention inside a tile cannot see the neighbouring tile. The downsampled whole-image block, prepended or appended to the tile blocks, is the only place the model can see the layout of the page as a whole. Without it, the model would be reading a stack of unconnected close-ups with no map.
4

Native dynamic resolution

Keep the aspect ratio, let the token count float

Native dynamic resolution, the scheme behind Qwen-VL and InternVL, skips the resizing and the tiling grid entirely. The image is patched at its own size and aspect ratio, so a tall, narrow page is handled as a tall, narrow page rather than being forced into a square, and the token count is area divided by patch area, adjusted by the merge factor: a bigger image produces more tokens, and the merge can pull that count back down. The encoder is modified to accept variable-length patch sequences, which means it can no longer assume a fixed grid of, say, $16 \times 16$ tokens. To keep the model oriented, these encoders usually use a two-dimensional rotary position embedding, a way of writing position into each token by rotating its vector by an angle that depends on the token's row and column. Position embeddings are the addresses that tell an otherwise order-blind transformer where each patch came from; a two-dimensional version encodes both coordinates at once, and because the address is a continuous function of row and column rather than a fixed lookup table, the model can generalise to grids it was never trained on. There is no thumbnail and no boundary cut, so nothing is lost at the seams.

What replaces the fixed budget is a decision per image: the pipeline picks a resolution, sometimes rounded up to the nearest multiple of the patch and merge sizes, based on the task and the context available. The rounding matters because the grid has to divide evenly into patches and then into merge blocks; a resolution that leaves a partial row would otherwise have to be dropped or padded. Document and OCR requests get many tokens; casual descriptions get few. OCR stands for optical character recognition, the task of reading text out of an image, and it is the most resolution-hungry request there is, because a single character may occupy only a handful of pixels. The trade is now made at request time rather than fixed once in training, which is the design's main advantage: the same model can spend a couple of hundred tokens on a casual photo and several thousand on a photographed contract. The plot below shows the native count rising with the square of the side, and the effect of the merge from the next step as a dashed line that grows much more slowly.

A second lever sits inside the encoder rather than at the token count. Qwen2.5-VL scales its vision transformer with windowed self-attention: most ViT layers attend only inside local $112 \times 112$ windows — a fixed neighbourhood of patches, so each patch sees its neighbours directly and nothing farther away — and full attention, the expensive kind that compares every patch with every other, runs at just a few layers. Because the quadratic term in attention grows with the square of the whole token grid, keeping full attention to a handful of layers means the encoder never pays that price over the whole grid at once, so the cost of a large image stops exploding as it scales up. This is a different answer from tiling, which cuts the image into pieces and accepts the seams, and from merging, which shrinks the sequence after the encoder has already paid for it; windowed attention leaves the token grid intact and changes what each layer is allowed to read. Where this shows up is Pixtral 12B, which handles several images in one sequence at their native resolution: each image keeps its own aspect ratio and its own token count rather than being forced into a common square, so a wide panorama, a tall page and a small photo can sit in the same prompt side by side, and the model sees several images of different shapes in one sequence rather than one image at a time.

Native token count against image side, with and without the merge, from Multimodal.tiles.nativeTokens. The dashed line marks the current input side.

Doubling the side quadruples the token count. That is the quadratic that the merge exists to soften.

5

Compressing the tokens

Four patches, one token

A merge layer groups a small block of neighbouring patch tokens into a single token before they reach the language model, which cuts the sequence by the square of the block size. This is the dicing-an-onion move in the other direction: rather than handing every grain of the dice to the language model, you pack a $2 \times 2$ cluster of grains into one larger parcel. A $2 \times 2$ merge divides the token count by four; a $4 \times 4$ merge by sixteen. A real model ships at exactly that number: InternVL3.5's dynamic-resolution path encodes each image patch as 1024 visual tokens and pixel-shuffles them to 256 before the language model sees them, the fourfold reduction the $2 \times 2$ block describes. Pooling does this by averaging the block, which is cheap and destroys fine detail. Averaging is a smoothing operation: if one patch carries a small, unique mark and its three neighbours carry background, the mean keeps the background and washes the mark out, and the information is gone before the language model ever sees it. Pixel-shuffle does it by concatenating the block's channels instead of averaging them: a $2 \times 2$ block of $C$-dimensional patches becomes one token of dimension $4C$, so no information is thrown away at the merge itself — the four patches' features are laid side by side into a wider vector, and the language model is asked to work with wider vectors instead of more of them. Perceiver-style resampling, which is the Q-Former's trick, is the learned version of the same idea, replacing the fixed block with learned queries that decide what to keep. A fixed merge treats every region of the image as equally important; a learned one can spend its tokens where the content actually is.

The bars below compare the options at the current input size. The token reduction is real and large, and it is the single most effective lever a practitioner has for fitting high-resolution input into a fixed context. A $4 \times 4$ merge cuts the sequence to a sixteenth of its length, which is the difference between a document that fits in the context window and one that does not. The caveat is the same one that applies to any compression: the detail that disappears at the merge cannot be recovered downstream. Once four patches have been averaged into one vector, no later layer can tell what the individual patches held, so the factor has to be chosen with the task in mind rather than by picking the largest one that fits.

Image tokens at the current input side for each pooling option, from Multimodal.cost.imageTokens. Pixel-shuffle matches pooling on count but keeps the block's channels.

A learned resampler sits at the same point in the pipeline; it just decides the compression from the image rather than by a fixed block.

💡 Pruning is the other half of compression. Rather than merging every block, an importance-pruning scheme scores patches and drops the ones that carry little information — blank background, uniform sky, whitespace between paragraphs. It can preserve more detail per token than a fixed merge, at the cost of a scoring pass and an irregular sequence length. The output is a different length for every image, which is the price of being selective.
6

Where this shows up

Tokens become FLOPs and bytes

The token count is the intermediate quantity; what a serving system feels is the prefill cost and the KV cache. Prefill runs the full attention and MLP stack over every prompt token at once, so its FLOP count grows with the sequence linearly in the projections and quadratically in attention. The MLP is the feed-forward block that transforms each token on its own, and the projections are the learned matrices that map a token into its query, key and value; both are per-token work, so they cost a fixed amount per token, whereas attention compares every token with every other and therefore grows with the square of the sequence. The KV cache then stores a key and a value for every token at every layer, so an image's tokens stay resident and are re-read on every decode step. That is the asymmetry to hold on to: prefill is a one-time arithmetic cost, while the cache is a permanent memory cost that is touched again for each new word the model writes. Both are computed here with Multimodal.cost.prefillFlops and Multimodal.cost.kvBytes.

The plots make the inversion that LLM Serving, Part 18 works out visible: a text-only request is often decode-dominated, but a request with a high-resolution image is a large prefill followed by cache-heavy decode, which is the opposite shape. A text-only prompt is short, so the model spends most of its time generating words one at a time, each step reading a small cache; a high-resolution image makes the prompt long, so the one big read at the start dominates, and every later step has to drag the whole image-shaped cache along with it. That is why multimodal serving tends to be memory-bandwidth bound — the bottleneck is how fast keys and values can be moved from memory, not how many FLOPs the arithmetic needs — and why compression at the vision side is worth more than it looks.

Prefill FLOPs and KV cache bytes against total prompt tokens, with the current image's contribution marked. The text prompt is assumed to be 64 tokens.

Head dimension is fixed at 128. Reducing KV heads with grouped-query attention cuts the cache in proportion without touching the FLOPs.

⚠ Compression has two bills, not one. A high-resolution image raises prefill FLOPs once and the KV cache forever after, so the cost of a generous token budget is paid on every generated token, not just the first. That is the practical reason a vision front end that outputs a fixed, short token set stays attractive even when a native-resolution encoder is available: shrinking that token set is a recurring saving rather than a one-time one.

Further reading

These papers trace the path from a fixed $224^2$ input to native dynamic resolution, and the serving analysis on the other side of the token budget. Read them in order and the design pressure becomes visible: each one is reacting to the resolution limit of the paper before it, and the serving reference is there to show what the accumulated token count costs once the model is actually deployed.

Cheat sheet

TermMeaning here
Fixed patchesResize to one square, patchify, constant token count; detail lost in the resize
Token count$(H/P)^2$ before merging; doubles the side, quadruples the tokens
AnyResEncode a grid of tiles plus a whole-image thumbnail; tokens scale with tile count
Grid pinpointsChoosing the tile layout whose aspect ratio best matches the image
Native resolutionPatch at the image's own size and aspect ratio, with 2-D position embeddings
2-D RoPERotary position encoding with separate row and column phases, so unseen grids are usable
Merge / poolingGroup a $k \times k$ block of patches into one token, dividing the budget by $k^2$
Pixel-shuffleRepack a block into channels instead of averaging, keeping detail at the same token count
PruningScore patches and drop the uninformative ones; irregular but detail-preserving
Prefill vs decodeAn image adds a large one-time prefill and a permanent KV cache
8

Check your understanding

0/4 answered