Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

Where the cost is

Encoder, prefill and decode are three different bills

A multimodal request is paid for in three stages, and it is worth separating them because each one is limited by a different resource. The encoder runs the vision tower over every patch, a compute-bound convolution-plus-attention pass over the whole image. Being compute-bound means the stage is limited by how fast the GPU can perform arithmetic rather than by how fast it can move data. The prefill then runs the language stack over the image tokens and the text prompt at once, which is compute-bound and carries the quadratic attention term. Prefill is the single first pass in which the model reads the whole prompt before it generates anything, and its attention term is quadratic because every token in that prompt compares itself with every other token, so doubling the prompt length roughly quadruples that part of the work. The decode finally generates the answer one token at a time, re-reading the whole KV cache on every step, which is bandwidth-bound. The KV cache is the running notebook of keys and values the model keeps for everything it has already read, so it does not redo the prompt for every new word, and decode is bandwidth-bound because it spends its time moving that notebook out of memory rather than doing arithmetic on it. The sizes are very different: an image of two thousand tokens is a large prefill and a permanently larger cache, so the request drifts from the decode-heavy shape a chat turn has toward a prefill-and-cache-heavy shape. That is the inversion an image causes, and Part 18 of the serving guide states it for a text workload.

The plot draws all three against image size, so you can watch the balance between them shift as the image grows. Encoder time is linear in the token count and divided by the encoder pool, where a pool is the set of identical GPU workers the stage is spread across; image encodes are independent of each other, so they parallelise cleanly across that pool. Prefill is roughly quadratic in the young sequence, because at that point the prompt is short and the quadratic attention term is still modest compared with the per-token projections. Decode grows with the cache that the image tokens leave behind, and that cache does not shrink once the prompt is read. The readout turns the three into a total and names the moment when the encoder is a large enough share of it to be worth pulling onto its own pool, which is the decision the next step makes properly. Until it is, the encoder shares the language pool and its cost hides inside the decode-heavy total.

Encoder, prefill and decode time against image tokens, from Multimodal.cost.prefillFlops and Multimodal.cost.kvBytes. The marker is the current image size.

💡 By the end of this part you'll be able to separate the encoder, prefill and decode bills for one request, explain why an image inverts the prefill-and-decode ratio, decide when to run the encoder on its own pool, quantify the saving from caching repeated embeddings, and name the open problems in evaluation and safety. Together those five skills are the whole serving picture: what a request costs, how to split it, how to avoid repeating work, and what the arithmetic leaves to judgement.
2

Disaggregating the encoder

A different bottleneck on a different pool

The encoder and the decoder want opposite things from a GPU, and that opposition is the whole reason to consider separating them. Batching means running many requests through the hardware at once so it is kept busy. The encoder processes a whole image at once, so it saturates compute and prefers large batches of images; the decoder processes one token at a time, so it is memory-bandwidth-bound and prefers a large batch of sequences. Memory-bandwidth-bound means the stage is limited by how quickly keys and values can be moved out of memory, not by how many arithmetic operations it can perform. Running both on the same devices means they interfere: an image encode arrives, takes the SMs — the streaming multiprocessors, the GPU's arithmetic units — for a few milliseconds, and stalls every decode step that was mid-flight, spiking the inter-token latency for every other user. Inter-token latency is the delay between one generated word and the next, and it is what a person actually feels when a reply types out slowly, so a single burst of image work can raise it for every conversation sharing the machine. This is the same argument that produced prefill-decode disaggregation in LLM Serving, Part 13, applied one stage earlier. Disaggregation here means running the two kinds of work on separate pools of devices instead of the same ones.

Disaggregating the encoder gives it its own pool, sized for image throughput, and ships the resulting embeddings to the language pool over the fabric. Throughput here means how many images the pool can encode per second, as opposed to how quickly any single one of them finishes, and the fabric is the high-speed interconnect, such as InfiniBand or NVLink, that joins the machines in a serving cluster. An embedding is the vector an image patch becomes after the encoder has processed it, so a two-thousand-token image is two thousand such vectors, each as wide as the model's hidden dimension. The trade is the transfer: embeddings are large — thousands of vectors of the model's width — so the link has to carry them faster than the encoder would have blocked the decoder for it to be a win. The bars compare a co-located request, where encoder and decoder share the same devices, with a disaggregated one at the current settings; the readout says whether the encoder is a large enough share of the bill to justify the split, and how many workers the pool has.

Total request time when the encoder shares the decode pool (it blocks) and when it runs on its own pool (it overlaps, at the cost of shipping embeddings).

The encoder is only worth disaggregating when it is a meaningful share of the request. For a small image on an idle GPU the transfer costs more than the block it removes.

⚠ Disaggregation moves the bottleneck, it does not remove it. A dedicated encoder pool sized for peak image traffic sits idle on a text-only hour, and a pool sized for the average queues images at the peak. As with prefill-decode splitting, the win is a goodput win under a mixed load, not a free speedup. Goodput counts the requests that complete within their target time rather than the raw rate, so the gain shows up as fewer stalls across all users, not as a shorter encode.
3

Caching embeddings

The same image should be encoded once

Many multimodal workloads send the same image again and again, and the repetition is often invisible from the vantage point of any single request. A document assistant re-uploads the same page as the conversation moves; a computer-use agent screenshots an interface that has not changed; a batch job runs a hundred questions against one chart. In all of those the encoder is doing identical work, and its output — the image's embedding vectors — depends only on the pixels and the encoder weights, not on the question being asked about them. An encoder cache exploits exactly that: caching those embeddings, keyed by a hash of the image so identical pixels map to the same entry, turns a repeated encode into a memory read, exactly as prefix caching turns a repeated prompt into a KV read in LLM Serving, Part 8. The saving is bounded by the hit rate — the fraction of requests whose image is already in the cache — which is a property of the workload rather than of the model, so the same cache is worth far more to a document assistant than to a camera that never sees the same photograph twice. The cost is the cache's memory: embeddings are image tokens times the model's width, which for a two-thousand-token image is tens of megabytes apiece.

The plot shows request time as the hit rate rises. At zero every request encodes; at one the encoder never runs for a repeat and only the prefill and decode remain, so the curve flattens at the cost of those two stages and cannot fall below it. The marker is the hit rate on the slider, and the readout reports the encoder time saved and the size of one cached image, which is the memory the cache spends per distinct picture it remembers.

Request time against encoder-cache hit rate at the current settings. The dashed line is the no-cache total; the curve is what caching removes in proportion to the hits.

Caching is a workload bet. Screenshots and documents repeat and cache well; one-shot photographs do not, and the cache is pure memory cost.

💡 Two caches, two keys. An encoder cache is keyed by the image and saves the vision tower; a prefix cache is keyed by the token sequence and saves the language model's KV. A multimodal request can hit both, and the two evolve on different timescales because the image changes far less often than the conversation does. Keeping them separate is also what makes each saving legible: a miss in one need not evict the other.
4

Safety

The pixels are an instruction channel too

Text alignment is trained on text, and a multimodal model accepts a second input channel that was never part of that training. Alignment here means teaching the model, from examples of harmful and harmless exchanges, which requests to decline. Two consequences follow, and both are architectural rather than incidental. First, the vision path is a new attack surface: an adversarial perturbation, a pattern of tiny pixel changes computed against the model and invisible to a person, can push a model past its refusal behaviour, because the perturbation is optimised in the same gradient space the language model reads. A refusal is the trained response in which the model declines a request instead of carrying it out, and defeating one this way is what the field calls a jailbreak. Second, and more directly exploitable, any text rendered inside an image is instructions to the model. A screenshot, a scanned form, a web page, a PDF, a caption burned into a video frame — all of them can carry a phrase that the model reads and obeys, with no user ever typing it.

Image jailbreaks

Pixels that flip a refusal

Optimised perturbations added to an image can suppress a model's safety training, so a request that would be refused in text succeeds when the same content is smuggled through the vision channel. The reason the attack transfers is that the perturbation rides in on the pixels but is computed against the model's own gradients, so it targets the same weights that the refusal behaviour was trained into. The defence is the same one text alignment uses — training on the attack, adding those perturbed images to the refusal examples — and it is a moving target, because the attack is re-optimised against each new defence and a defence that has not seen the latest perturbation is unlikely to hold.

Prompt injection

Text in an image is untrusted input

This failure is called indirect prompt injection: a document or a screenshot can contain a sentence addressed to the model, and the model has no reliable way to tell it from the user's instruction, because to the language model both arrive as ordinary tokens in the same context window. When the model can also ground and click, the injected instruction becomes an action rather than a topic: an adversarial page can ask a computer-use agent to exfiltrate data or open a file. To exfiltrate is to move data out of the system to somewhere the attacker can read it. Untrusted images must be treated as untrusted code.

⚠ Grounding converts a description into a capability. A model that can only caption an injected instruction is a nuisance; a model that can ground it to a button and click is an agent acting on someone else's sentence. The permission boundary, not the model's refusal training, is what has to hold, which shifts the question from what the model can be persuaded to say to what the surrounding system allows it to touch.
5

Open problems

What the arithmetic does not settle

The engineering questions in this part have answers, which is why the cost model, the split and the cache can all be written down and tuned. The questions below do not, and they are the ones a practitioner should carry out of the volume. They are open in the specific sense that a plausible design choice exists for each and no benchmark yet decides between them. The six below are the ones this volume leaves unsettled: contaminated evaluation, the understanding-versus-generation settlement, grounding precision, resolution against cost, end-to-end evaluation of agents, and safety for a grounded model. That makes them worth knowing precisely because they are the places where a confident opinion is more likely to be a preference than a result.

Evaluation that is not contaminated

MMMU, MathVista, DocVQA, ChartQA and MMBench are the benchmarks — standardised test sets — behind the scores quoted in model cards, and each has a way of leaking. A benchmark leaks when something other than the ability it names can produce a good score: knowledge-heavy items answerable from the question text alone, template charts shared between train and test, documents solvable by OCR alone, and judges whose bias is part of the score. A high number with no held-out set is weak evidence: held-out means the test items were deliberately kept out of training, so a model that has never seen them cannot be scoring from memory. Freshly collected, dynamically generated, or adversarially held-out evaluations are the direction, and none is standard yet.

The understanding and generation settlement

The objective conflict from the unified-model part is unresolved. That conflict is between training one model to understand images and to generate them, when the two directions pull on the shared representation in different ways. Separate heads, continuous latent heads, and mixture-of-experts trunks are all live positions: a head is the small output layer that turns the trunk's shared representation into predictions, and a mixture-of-experts trunk splits the network into many expert subnetworks and routes each token to a few of them, so different parts can specialise. There is no clean benchmark that says which representation best serves both directions at once. Every unified model is a point on a line nobody has measured the whole of.

Grounding precision and reliability

Coordinate tokens set a hard precision floor, because a location written as a discrete token can only land on a smallest rectangle the vocabulary can express. On top of that floor, the boxes and masks a model emits are not consistently calibrated: the same query returns a tight box on one image and a loose one on the next, so the reported region cannot be trusted to the same degree twice. A mask is the finer-grained alternative to a box, an outline of the object at pixel level rather than a rectangle around it. For an interface agent a few pixels is the difference between a click and a miss, and there is no standard method for bounding that error the way a latency SLO — a service-level objective, the agreed target for response time — bounds a serving system.

Resolution against cost

Native-resolution encoders keep the detail that documents and charts need, and every token of it is paid for at prefill and again on every decode step. Native resolution means patching the image at its own size instead of shrinking it to a fixed square, and the merge and pruning tricks of the token-budget part are the two ways to buy back room: merging groups neighbouring patches into one token, while pruning drops the uninformative ones outright. Either way the right point on that trade depends on the task and the context in a way no single architecture has settled, since a photographed contract and a casual snapshot are the same kind of input to the model but want very different token budgets. FastV is the complementary move on the inference side: it drops a large fraction of image tokens after an early language-model layer, ranked by the attention they actually receive, so the saving is made at serve time on the tokens the prompt produced rather than at train time on every input the way the merge does.

End-to-end evaluation of agents

A VLA — a vision-language-action model, one that turns what it sees and reads into motor commands — or a computer-use agent is graded on task success, which depends on the environment, the controller and the retries as much as on the model. That is the difficulty: the score is not a property of the model alone, so a change in the lab, the robot, or the website can move it without any change to the model. Reproducing a robot result on new hardware, or a web-agent result on a changed site, is still harder than it should be, so cross-system comparison is weak.

Safety for a grounded model

Refusal training assumes the dangerous content arrives as a request. A multimodal agent breaks that assumption twice: it reads instructions from images, where the text is content rather than a turn in a conversation, and it acts on them, so declining to say something is no longer the only thing at stake. The relevant boundary is therefore the set of capabilities the system is allowed to exercise, not the set of phrases the model is trained to decline.

Further reading

These sources cover the serving split this part reuses, the caching it borrows, and the two safety failures that are specific to a channel that carries instructions as pixels. Read the first two together and the pattern is visible: the same argument that justified splitting prefill from decode justifies separating the encoder one stage earlier, and the same idea of a reusable cache applies to a repeated image as it does to a repeated prompt prefix.

Cheat sheet

TermMeaning here
Encoder stageRunning the vision tower over the patches; compute-bound and divisible across a pool
PrefillRunning the language stack over image plus text tokens at once; carries the quadratic term
DecodeOne token at a time, re-reading the enlarged KV cache; bandwidth-bound
The inversionAn image makes a request prefill- and cache-heavy where a chat turn is decode-heavy
Encoder disaggregationIts own pool, sized for image throughput, with embeddings shipped to the language pool
Embedding cacheReusing a vision tower's output for a repeated image, keyed by the picture
Prefix cacheReusing the language model's KV for a repeated token prefix; the other half of the pair
Image jailbreakA pixel perturbation that suppresses refusal behaviour trained on text
Indirect injectionAn instruction a model reads from content rather than from the user; a screenshot is content
Open problemsContamination-proof evaluation, the objective settlement, grounding calibration, cost against resolution, agent evaluation, and safety for a grounded model
7

Check your understanding

0/4 answered