Serving and evaluating multimodal systems
A text request is a stream of tokens that arrives and is answered. A multimodal request — one that carries an image alongside its text — has a second workload hiding inside it: before the language model sees a single image token, a vision encoder has to turn the pixels into them. The encoder is the vision tower from earlier in this volume: the half of the system that cuts an image into patches and turns each patch into a vector of numbers the language model can read. That encoder runs different kernels — different low-level GPU programs than the language model uses — has a different bottleneck, and can be cached independently. A bottleneck is whichever resource a stage is waiting on, and the two stages wait on different things: the encoder is limited by arithmetic, the decoding half by memory traffic, which is why the serving story for a vision-language model is not the language-model story with a longer prompt. This part prices the request, splits the encoder from the decoder, caches the embeddings that repeat (the vectors the encoder outputs, which depend only on the picture), and then turns to the two things the arithmetic cannot settle: safety, and the open problems the field has not closed. The first half is engineering, where a measurement settles the question; the second half is judgement, where it does not.
Where the cost is
Encoder, prefill and decode are three different bills
A multimodal request is paid for in three stages, and it is worth separating them because each one is limited by a different resource. The encoder runs the vision tower over every patch, a compute-bound convolution-plus-attention pass over the whole image. Being compute-bound means the stage is limited by how fast the GPU can perform arithmetic rather than by how fast it can move data. The prefill then runs the language stack over the image tokens and the text prompt at once, which is compute-bound and carries the quadratic attention term. Prefill is the single first pass in which the model reads the whole prompt before it generates anything, and its attention term is quadratic because every token in that prompt compares itself with every other token, so doubling the prompt length roughly quadruples that part of the work. The decode finally generates the answer one token at a time, re-reading the whole KV cache on every step, which is bandwidth-bound. The KV cache is the running notebook of keys and values the model keeps for everything it has already read, so it does not redo the prompt for every new word, and decode is bandwidth-bound because it spends its time moving that notebook out of memory rather than doing arithmetic on it. The sizes are very different: an image of two thousand tokens is a large prefill and a permanently larger cache, so the request drifts from the decode-heavy shape a chat turn has toward a prefill-and-cache-heavy shape. That is the inversion an image causes, and Part 18 of the serving guide states it for a text workload.
The plot draws all three against image size, so you can watch the balance between them shift as the image grows. Encoder time is linear in the token count and divided by the encoder pool, where a pool is the set of identical GPU workers the stage is spread across; image encodes are independent of each other, so they parallelise cleanly across that pool. Prefill is roughly quadratic in the young sequence, because at that point the prompt is short and the quadratic attention term is still modest compared with the per-token projections. Decode grows with the cache that the image tokens leave behind, and that cache does not shrink once the prompt is read. The readout turns the three into a total and names the moment when the encoder is a large enough share of it to be worth pulling onto its own pool, which is the decision the next step makes properly. Until it is, the encoder shares the language pool and its cost hides inside the decode-heavy total.
Encoder, prefill and decode time against image tokens, from Multimodal.cost.prefillFlops and Multimodal.cost.kvBytes. The marker is the current image size.
Disaggregating the encoder
A different bottleneck on a different pool
The encoder and the decoder want opposite things from a GPU, and that opposition is the whole reason to consider separating them. Batching means running many requests through the hardware at once so it is kept busy. The encoder processes a whole image at once, so it saturates compute and prefers large batches of images; the decoder processes one token at a time, so it is memory-bandwidth-bound and prefers a large batch of sequences. Memory-bandwidth-bound means the stage is limited by how quickly keys and values can be moved out of memory, not by how many arithmetic operations it can perform. Running both on the same devices means they interfere: an image encode arrives, takes the SMs — the streaming multiprocessors, the GPU's arithmetic units — for a few milliseconds, and stalls every decode step that was mid-flight, spiking the inter-token latency for every other user. Inter-token latency is the delay between one generated word and the next, and it is what a person actually feels when a reply types out slowly, so a single burst of image work can raise it for every conversation sharing the machine. This is the same argument that produced prefill-decode disaggregation in LLM Serving, Part 13, applied one stage earlier. Disaggregation here means running the two kinds of work on separate pools of devices instead of the same ones.
Disaggregating the encoder gives it its own pool, sized for image throughput, and ships the resulting embeddings to the language pool over the fabric. Throughput here means how many images the pool can encode per second, as opposed to how quickly any single one of them finishes, and the fabric is the high-speed interconnect, such as InfiniBand or NVLink, that joins the machines in a serving cluster. An embedding is the vector an image patch becomes after the encoder has processed it, so a two-thousand-token image is two thousand such vectors, each as wide as the model's hidden dimension. The trade is the transfer: embeddings are large — thousands of vectors of the model's width — so the link has to carry them faster than the encoder would have blocked the decoder for it to be a win. The bars compare a co-located request, where encoder and decoder share the same devices, with a disaggregated one at the current settings; the readout says whether the encoder is a large enough share of the bill to justify the split, and how many workers the pool has.
Total request time when the encoder shares the decode pool (it blocks) and when it runs on its own pool (it overlaps, at the cost of shipping embeddings).
The encoder is only worth disaggregating when it is a meaningful share of the request. For a small image on an idle GPU the transfer costs more than the block it removes.
Caching embeddings
The same image should be encoded once
Many multimodal workloads send the same image again and again, and the repetition is often invisible from the vantage point of any single request. A document assistant re-uploads the same page as the conversation moves; a computer-use agent screenshots an interface that has not changed; a batch job runs a hundred questions against one chart. In all of those the encoder is doing identical work, and its output — the image's embedding vectors — depends only on the pixels and the encoder weights, not on the question being asked about them. An encoder cache exploits exactly that: caching those embeddings, keyed by a hash of the image so identical pixels map to the same entry, turns a repeated encode into a memory read, exactly as prefix caching turns a repeated prompt into a KV read in LLM Serving, Part 8. The saving is bounded by the hit rate — the fraction of requests whose image is already in the cache — which is a property of the workload rather than of the model, so the same cache is worth far more to a document assistant than to a camera that never sees the same photograph twice. The cost is the cache's memory: embeddings are image tokens times the model's width, which for a two-thousand-token image is tens of megabytes apiece.
The plot shows request time as the hit rate rises. At zero every request encodes; at one the encoder never runs for a repeat and only the prefill and decode remain, so the curve flattens at the cost of those two stages and cannot fall below it. The marker is the hit rate on the slider, and the readout reports the encoder time saved and the size of one cached image, which is the memory the cache spends per distinct picture it remembers.
Request time against encoder-cache hit rate at the current settings. The dashed line is the no-cache total; the curve is what caching removes in proportion to the hits.
Caching is a workload bet. Screenshots and documents repeat and cache well; one-shot photographs do not, and the cache is pure memory cost.
Safety
The pixels are an instruction channel too
Text alignment is trained on text, and a multimodal model accepts a second input channel that was never part of that training. Alignment here means teaching the model, from examples of harmful and harmless exchanges, which requests to decline. Two consequences follow, and both are architectural rather than incidental. First, the vision path is a new attack surface: an adversarial perturbation, a pattern of tiny pixel changes computed against the model and invisible to a person, can push a model past its refusal behaviour, because the perturbation is optimised in the same gradient space the language model reads. A refusal is the trained response in which the model declines a request instead of carrying it out, and defeating one this way is what the field calls a jailbreak. Second, and more directly exploitable, any text rendered inside an image is instructions to the model. A screenshot, a scanned form, a web page, a PDF, a caption burned into a video frame — all of them can carry a phrase that the model reads and obeys, with no user ever typing it.
Pixels that flip a refusal
Optimised perturbations added to an image can suppress a model's safety training, so a request that would be refused in text succeeds when the same content is smuggled through the vision channel. The reason the attack transfers is that the perturbation rides in on the pixels but is computed against the model's own gradients, so it targets the same weights that the refusal behaviour was trained into. The defence is the same one text alignment uses — training on the attack, adding those perturbed images to the refusal examples — and it is a moving target, because the attack is re-optimised against each new defence and a defence that has not seen the latest perturbation is unlikely to hold.
Text in an image is untrusted input
This failure is called indirect prompt injection: a document or a screenshot can contain a sentence addressed to the model, and the model has no reliable way to tell it from the user's instruction, because to the language model both arrive as ordinary tokens in the same context window. When the model can also ground and click, the injected instruction becomes an action rather than a topic: an adversarial page can ask a computer-use agent to exfiltrate data or open a file. To exfiltrate is to move data out of the system to somewhere the attacker can read it. Untrusted images must be treated as untrusted code.
Open problems
What the arithmetic does not settle
The engineering questions in this part have answers, which is why the cost model, the split and the cache can all be written down and tuned. The questions below do not, and they are the ones a practitioner should carry out of the volume. They are open in the specific sense that a plausible design choice exists for each and no benchmark yet decides between them. The six below are the ones this volume leaves unsettled: contaminated evaluation, the understanding-versus-generation settlement, grounding precision, resolution against cost, end-to-end evaluation of agents, and safety for a grounded model. That makes them worth knowing precisely because they are the places where a confident opinion is more likely to be a preference than a result.
Evaluation that is not contaminated
MMMU, MathVista, DocVQA, ChartQA and MMBench are the benchmarks — standardised test sets — behind the scores quoted in model cards, and each has a way of leaking. A benchmark leaks when something other than the ability it names can produce a good score: knowledge-heavy items answerable from the question text alone, template charts shared between train and test, documents solvable by OCR alone, and judges whose bias is part of the score. A high number with no held-out set is weak evidence: held-out means the test items were deliberately kept out of training, so a model that has never seen them cannot be scoring from memory. Freshly collected, dynamically generated, or adversarially held-out evaluations are the direction, and none is standard yet.
The understanding and generation settlement
The objective conflict from the unified-model part is unresolved. That conflict is between training one model to understand images and to generate them, when the two directions pull on the shared representation in different ways. Separate heads, continuous latent heads, and mixture-of-experts trunks are all live positions: a head is the small output layer that turns the trunk's shared representation into predictions, and a mixture-of-experts trunk splits the network into many expert subnetworks and routes each token to a few of them, so different parts can specialise. There is no clean benchmark that says which representation best serves both directions at once. Every unified model is a point on a line nobody has measured the whole of.
Grounding precision and reliability
Coordinate tokens set a hard precision floor, because a location written as a discrete token can only land on a smallest rectangle the vocabulary can express. On top of that floor, the boxes and masks a model emits are not consistently calibrated: the same query returns a tight box on one image and a loose one on the next, so the reported region cannot be trusted to the same degree twice. A mask is the finer-grained alternative to a box, an outline of the object at pixel level rather than a rectangle around it. For an interface agent a few pixels is the difference between a click and a miss, and there is no standard method for bounding that error the way a latency SLO — a service-level objective, the agreed target for response time — bounds a serving system.
Resolution against cost
Native-resolution encoders keep the detail that documents and charts need, and every token of it is paid for at prefill and again on every decode step. Native resolution means patching the image at its own size instead of shrinking it to a fixed square, and the merge and pruning tricks of the token-budget part are the two ways to buy back room: merging groups neighbouring patches into one token, while pruning drops the uninformative ones outright. Either way the right point on that trade depends on the task and the context in a way no single architecture has settled, since a photographed contract and a casual snapshot are the same kind of input to the model but want very different token budgets. FastV is the complementary move on the inference side: it drops a large fraction of image tokens after an early language-model layer, ranked by the attention they actually receive, so the saving is made at serve time on the tokens the prompt produced rather than at train time on every input the way the merge does.
End-to-end evaluation of agents
A VLA — a vision-language-action model, one that turns what it sees and reads into motor commands — or a computer-use agent is graded on task success, which depends on the environment, the controller and the retries as much as on the model. That is the difficulty: the score is not a property of the model alone, so a change in the lab, the robot, or the website can move it without any change to the model. Reproducing a robot result on new hardware, or a web-agent result on a changed site, is still harder than it should be, so cross-system comparison is weak.
Safety for a grounded model
Refusal training assumes the dangerous content arrives as a request. A multimodal agent breaks that assumption twice: it reads instructions from images, where the text is content rather than a turn in a conversation, and it acts on them, so declining to say something is no longer the only thing at stake. The relevant boundary is therefore the set of capabilities the system is allowed to exercise, not the set of phrases the model is trained to decline.
Further reading
These sources cover the serving split this part reuses, the caching it borrows, and the two safety failures that are specific to a channel that carries instructions as pixels. Read the first two together and the pattern is visible: the same argument that justified splitting prefill from decode justifies separating the encoder one stage earlier, and the same idea of a reusable cache applies to a repeated image as it does to a repeated prompt prefix.
- Yinmin Zhong, Shengyu Liu, Junda Chen and coauthors, "DistServe: Disaggregating Prefill and Decoding for Goodput-optimized LLM Serving", OSDI 2024 — the prefill-decode split that encoder disaggregation extends one stage earlier.
- Ruoyu Qin, Zheming Li, Weiran He and coauthors, "Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving", 2024 — the KVCache store, the prefix-cache counterpart to the embedding cache here.
- Xiangyu Qi, Kaixiong Zhou, Yiming Li and coauthors, "Visual Adversarial Examples Jailbreak Aligned Large Language Models", 2023 — a perturbation optimised in pixel space that defeats text-only alignment.
- Kai Greshake, Sahar Abdelnabi, Shailesh Mishra and coauthors, "Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection", 2023 — instructions hidden in retrieved content, the mechanism a screenshot injection reuses.
- LLM Serving, Part 13: Disaggregation and Part 18: Workloads — the pooling and the prefill/decode accounting this part applies to vision.
Cheat sheet
| Term | Meaning here |
|---|---|
| Encoder stage | Running the vision tower over the patches; compute-bound and divisible across a pool |
| Prefill | Running the language stack over image plus text tokens at once; carries the quadratic term |
| Decode | One token at a time, re-reading the enlarged KV cache; bandwidth-bound |
| The inversion | An image makes a request prefill- and cache-heavy where a chat turn is decode-heavy |
| Encoder disaggregation | Its own pool, sized for image throughput, with embeddings shipped to the language pool |
| Embedding cache | Reusing a vision tower's output for a repeated image, keyed by the picture |
| Prefix cache | Reusing the language model's KV for a repeated token prefix; the other half of the pair |
| Image jailbreak | A pixel perturbation that suppresses refusal behaviour trained on text |
| Indirect injection | An instruction a model reads from content rather than from the user; a screenshot is content |
| Open problems | Contamination-proof evaluation, the objective settlement, grounding calibration, cost against resolution, agent evaluation, and safety for a grounded model |