Glossary & numbers to know
This is the multimodal guide's back matter: the reference material you keep open beside the parts, rather than a part you read in order. It holds one alphabetical index that links every term to the part that introduces it, plus the three tables the rest of the series leans on — the architecture eras from two-tower to omni, the token budget each resolution scheme actually spends, and the evaluation benchmarks with their contamination caveats. A contamination caveat is a warning that a benchmark's test items, or things very close to them, may have appeared in a model's training data, which makes a high score easier to earn and harder to trust. It is the companion to the serving guide's glossary, which shares the prefill/decode vocabulary this volume borrows (prefill reads the prompt in one pass; decode generates the answer one token at a time), and to the training guide's glossary for the modelling terms underneath.
Term index
Every term, linked to the part that introduces it
| Term | What it means | Introduced in |
|---|---|---|
| A | ||
| Action chunk | A horizon of actions emitted at once and executed open-loop, so one inference covers many control steps. | Part 15 |
| Action token | A discretised action dimension emitted as an ordinary token; a seven-axis arm emits seven per step. | Part 15 |
| AnyRes | Encode a grid of encoder-sized tiles plus a downsampled whole-image thumbnail; tokens scale with tile count. | Part 6 |
| Attention sink | A few positions that absorb a large share of attention while carrying little information, starving the informative image tokens. | Part 8 |
| B | ||
| Bridge | The projector, resampler or cross-attention interface that joins a frozen encoder to a language model; what early fusion removes. | Part 5 |
| C | ||
| Class token | An extra learned token whose final output is the whole-image summary a classifier head reads. | Part 1 |
| Codebook collapse | A learned VQ codebook concentrating onto a few entries, leaving most of its capacity dead. | Part 13 |
| Coordinate token | A quantised bin index standing in for a continuous coordinate; a box is four of them. | Part 14 |
| Cross-modal attention | Attention weights off the modality diagonal: text attending to patches, speech attending to text. | Part 9 |
| D | ||
| Data mixture | The text, image and instruction blend that fixes a vision-language model's language prior. | Part 7 |
| Decode | Generating one token at a time, re-reading the whole KV cache each step; memory-bandwidth-bound. | Part 16 |
| Diffusion head | A small denoiser on the trunk's image-latent positions that regresses noise or velocity instead of emitting a token id. | Part 13 |
| E | ||
| Early fusion | One trunk pretrained on interleaved tokens from every modality, with per-modality embedders and no separate towers. | Part 9 |
| Embedding cache | Reusing a vision tower's output for a repeated image, keyed by a hash of the picture. | Part 16 |
| Encoder | The vision-tower stage that turns pixels into tokens; compute-bound and divisible across its own pool. | Part 16 |
| Encoder disaggregation | Running the vision encoder on a dedicated pool so an image encode cannot stall in-flight decode steps. | Part 16 |
| F | ||
| Flow matching | Regress the constant velocity along a straight path from noise to sample; sample with a few Euler steps. | Part 15 |
| FSQ | Finite scalar quantization: round each latent coordinate on a fixed grid, with no codebook to collapse. | Part 13 |
| G | ||
| Grounding | Making a model emit spatial coordinates — boxes, points, masks — as tokens, after the phrase that names the object. | Part 14 |
| GUI grounding | Localising interface elements to boxes so a computer-use agent can emit a click or keystroke at them. | Part 14 |
| H | ||
| Hallucination | Naming or describing something the pixels do not support, because the language prior outweighs the image evidence. | Part 8 |
| I | ||
| Image jailbreak | An optimised pixel perturbation that suppresses refusal behaviour a model learned from text. | Part 16 |
| Image tokens | Vision input expanded into context tokens; large, prefill-heavy and a cause of the prefill/decode inversion. | Part 6 |
| Inductive bias | Locality and translation equivariance: a convolution has them, a patch transformer learns them from data. | Part 1 |
| InfoNCE | The symmetric contrastive loss on a similarity matrix that trains CLIP-style two-tower models. | Part 2 |
| K | ||
| KV cache (multimodal) | The keys and values an image's tokens add; resident for the request and re-read on every decode step. | Part 16 |
| L | ||
| Latent world model | Predicts the next observation's representation so a policy can roll a plan forward in imagination. | Part 15 |
| M | ||
| Masked generation (MAR) | Reveal positions in a seeded order and predict masked positions in parallel, with context on both sides. | Part 13 |
| Merge (2×2) | Group a block of four patches into one token before the language model, dividing the token budget by four. | Part 6 |
| Modality embedder | The per-modality front end that maps a tokenizer's output into the trunk's width. | Part 9 |
| Modality-aware experts | A shared attention stack with a routed feed-forward pool that can specialise per modality. | Part 9 |
| N | ||
| Native dynamic resolution | Patch at the image's own size and aspect ratio, with 2-D position embeddings and a floating token count. | Part 6 |
| Next-scale (VAR) | Predict whole scales coarse to fine, so sequential depth grows logarithmically rather than quadratically. | Part 13 |
| P | ||
| Patchify | Cut an image into patches and flatten each into a vector of numbers the transformer can read. | Part 1 |
| Pixel-shuffle | Repack a patch block into channels instead of averaging it, keeping detail at the same token count. | Part 6 |
| Position embedding | Added to each patch so permutation-invariant attention knows where a patch came from; learned or sinusoidal. | Part 1 |
| Prefill | Running the language stack over image plus text tokens at once; compute-bound and carries the quadratic term. | Part 16 |
| Prompt injection | An instruction a model reads from content rather than from the user; a screenshot is content. | Part 16 |
| Promptable segmentation | One trained model that returns a mask for a point, box or mask prompt without being fine-tuned per object. | Part 14 |
| Projector | The adapter that maps encoder vectors into the language model's embedding width; trained while the towers stay frozen. | Part 5 |
| Q | ||
| Q-Former | A querying bridge of a fixed set of learned queries between a frozen encoder and a language model. | Part 5 |
| Quantisation error | Half a bin for a coordinate token; the precision floor no model can point finer than. | Part 14 |
| R | ||
| Raster order | Row-major autoregression over image tokens; nearest neighbours in space are far apart in sequence. | Part 13 |
| Referring expression | A phrase that selects one object by describing it and its relations to other objects. | Part 14 |
| S | ||
| Shared embedding space | One vector space where a picture and its caption land near each other, enabling retrieval and zero-shot tasks. | Part 3 |
| SigLIP | A pairwise sigmoid contrastive objective, an alternative to the softmax InfoNCE loss. | Part 2 |
| Staged training | The freezing schedule and stage order — align, then fine-tune — that trains a vision-language model. | Part 7 |
| T | ||
| Temperature (contrastive) | The scale that sharpens or flattens the softmax in contrastive training and in attention. | Part 2 |
| Token budget | The number of image tokens a resolution scheme spends, and therefore its prefill and KV cost. | Part 6 |
| Token count | $(H/P)^2$ for a square image; the sequence length the transformer actually sees. | Part 1 |
| U | ||
| Unified model | One trunk that both reads images to text and generates images from text. | Part 13 |
| V | ||
| VLA | A vision-language-action model: a multimodal trunk whose output head emits actions, discrete or continuous. | Part 15 |
| VQ-VAE | An autoencoder whose encoder vectors snap to a learned codebook of discrete entries. | Part 13 |
| Zero-shot classification | Scoring an image against class-text embeddings in a shared space, with no classifier trained on those classes. | Part 3 |
Architecture eras
From two towers to one omni trunk
The field's five stages are not replacements so much as a widening of what the interface between vision and language is allowed to be. Fusion here means where the two modalities first meet: in a thin bridge between two separate towers, in every attention layer of one shared trunk, or somewhere in between. Read the table by asking what each stage actually trains. The earlier the fusion, the more of the model is in the loop, meaning the more components receive gradients and can adapt to one another, and the more of the model has to be pretrained jointly. Early stages reuse strong pretrained halves and therefore train comparatively little; later stages pay for a joint run and buy direct cross-modal mixing in return.
| Stage | Representative models | What it trains | What it leaves unsolved |
|---|---|---|---|
| Two-tower contrastive | CLIP, SigLIP | Both towers, with a contrastive loss into one shared embedding space | No generation, no instruction following; a similarity score, not an answer |
| Connector / Q-Former | BLIP-2, Flamingo, InstructBLIP | A small bridge over frozen encoders and a frozen language model | The interface is the only thing that learns; the encoder never hears what the LLM needs |
| Adapter / projector | LLaVA, Qwen-VL, InternVL | An MLP projector, often with the language model, on a mostly frozen encoder | Resolution via tiling; the encoder stays fixed at whatever it was trained for |
| Early fusion | Chameleon, Gemini, Fuyu | One trunk on interleaved text, image and audio tokens from the start | Needs a joint pretraining run; one shared, zero-sum token budget |
| Omni | GPT-4o, Gemini, Qwen2.5-Omni, Moshi | Any-to-any across text, image, audio and video, often with speech output | Real-time streaming and full-duplex turn-taking remain hard, and the token budget is shared by all |
Token budget cheat card
What each resolution scheme spends
An image is as many tokens as its patch grid says — the grid being the array of patch positions the encoder cut it into — and every scheme in the table is a different answer to the same question: keep the detail or keep the sequence short. The counts below use a 14-pixel patch, which is the standard CLIP/SigLIP geometry, and a 256-token per-tile block with a thumbnail for AnyRes, the thumbnail being a downsampled whole-image view that preserves global context while the tiles hold the detail. The arithmetic comes from Multimodal.tiles.nativeTokens, Multimodal.tiles.anyResTokens and Multimodal.cost.imageTokens; the FLOPs and bytes they imply — floating-point operations and memory traffic, the two quantities a serving bill is actually made of — are worked out in the resolution part.
| Scheme | Input | Tokens | Trade |
|---|---|---|---|
| Fixed 224² | 224 × 224, P = 14 | 256 | Constant cost, trivial batching; detail is lost in the resize |
| AnyRes 672² | 672 × 672, 336-px tiles | 1280 | 4 tiles + 1 thumbnail × 256; hard seams and a scaling cost |
| Native dynamic 448² | 448 × 448, P = 14 | 1024 | Aspect ratio kept, no seams; count floats with size |
| Native dynamic 896² | 896 × 896, P = 14 | 4096 | Doubling the side quadruples the tokens; the quadratic bites |
| 2×2 merge on 448² | 448 × 448, P = 14, merge 2 | 256 | Divide by four; pooling loses detail, pixel-shuffle keeps it in channels |
Evaluation benchmarks
What each measures, and how each leaks
The scores quoted in model cards — the short summary documents released alongside a model — come from a handful of benchmarks, where a benchmark is a fixed test set with a scoring rule. Each one has a specific way of being easier than the ability it names: some items can be guessed from the question alone, and some have leaked into training data. A high number with no contamination analysis is weaker evidence than a modest number with one, because the analysis is the only part that checks whether the test was genuinely unseen. Treat the middle column as the claim and the caveat column as the reason not to take it at face value.
| Benchmark | What it measures | Contamination and leakage caveats |
|---|---|---|
| MMMU | College-level reasoning across diagrams, charts, tables and figures | Knowledge-heavy; many items are answerable from the question text and subject knowledge without the image |
| MathVista | Mathematical reasoning grounded in visual contexts | A fraction of items are solvable from text alone; contamination from public solution text is hard to rule out |
| DocVQA | Question answering over scanned documents | Mostly an OCR test; needs high resolution and is saturating on leading models |
| ChartQA | Reasoning about values and trends in charts | Template-generated charts and augmented training data create leakage between train and test |
| MMBench | Many abilities at once, scored by a language-model judge | Judge bias and answer-style sensitivity: the score moves with the judge as well as with the model |
Further reading
Primary sources behind the numbers
- Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov and coauthors, "An Image is Worth 16x16 Words", 2020 — the patch transformer and the class token behind Part 1.
- Haotian Liu, Chunyuan Li, Yuheng Li and Yong Jae Lee, "Improved Baselines with Visual Instruction Tuning", 2023 — LLaVA and the AnyRes tiling of Part 6.
- Chameleon Team, "Chameleon: Mixed-Modal Early-Fusion Foundation Models", 2024 — the fused trunk of the architecture table.
- Keyu Tian, Yi Jiang, Zehuan Yuan and coauthors, "Visual Autoregressive Modeling", 2024 — next-scale generation, one of the three orders in Part 13.
- Alexander Kirillov, Eric Mintun, Nikhila Ravi and coauthors, "Segment Anything", 2023 — promptable segmentation in Part 14.
- Kevin Black, Noah Brown, Danny Driess and coauthors, "π0: A Vision-Language-Action Flow Model for General Robot Control", 2024 — the flow-matching action expert of Part 15.
- Yinmin Zhong, Shengyu Liu, Junda Chen and coauthors, "DistServe: Disaggregating Prefill and Decoding for Goodput-optimized LLM Serving", OSDI 2024 — the split Part 16 extends to the encoder.
- Xiang Yue, Yuansheng Ni, Kai Zhang and coauthors, "MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark", 2023 — the benchmark whose text-prior caveat the evaluation table records.