Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

Term index

Every term, linked to the part that introduces it

💡 Filter the list. Type any fragment — a term, a technology, a failure mode — and the table hides non-matching rows and reports how many remain. Matching is case-insensitive and looks at both the term and its definition.
TermWhat it meansIntroduced in
A
Action chunkA horizon of actions emitted at once and executed open-loop, so one inference covers many control steps.Part 15
Action tokenA discretised action dimension emitted as an ordinary token; a seven-axis arm emits seven per step.Part 15
AnyResEncode a grid of encoder-sized tiles plus a downsampled whole-image thumbnail; tokens scale with tile count.Part 6
Attention sinkA few positions that absorb a large share of attention while carrying little information, starving the informative image tokens.Part 8
B
BridgeThe projector, resampler or cross-attention interface that joins a frozen encoder to a language model; what early fusion removes.Part 5
C
Class tokenAn extra learned token whose final output is the whole-image summary a classifier head reads.Part 1
Codebook collapseA learned VQ codebook concentrating onto a few entries, leaving most of its capacity dead.Part 13
Coordinate tokenA quantised bin index standing in for a continuous coordinate; a box is four of them.Part 14
Cross-modal attentionAttention weights off the modality diagonal: text attending to patches, speech attending to text.Part 9
D
Data mixtureThe text, image and instruction blend that fixes a vision-language model's language prior.Part 7
DecodeGenerating one token at a time, re-reading the whole KV cache each step; memory-bandwidth-bound.Part 16
Diffusion headA small denoiser on the trunk's image-latent positions that regresses noise or velocity instead of emitting a token id.Part 13
E
Early fusionOne trunk pretrained on interleaved tokens from every modality, with per-modality embedders and no separate towers.Part 9
Embedding cacheReusing a vision tower's output for a repeated image, keyed by a hash of the picture.Part 16
EncoderThe vision-tower stage that turns pixels into tokens; compute-bound and divisible across its own pool.Part 16
Encoder disaggregationRunning the vision encoder on a dedicated pool so an image encode cannot stall in-flight decode steps.Part 16
F
Flow matchingRegress the constant velocity along a straight path from noise to sample; sample with a few Euler steps.Part 15
FSQFinite scalar quantization: round each latent coordinate on a fixed grid, with no codebook to collapse.Part 13
G
GroundingMaking a model emit spatial coordinates — boxes, points, masks — as tokens, after the phrase that names the object.Part 14
GUI groundingLocalising interface elements to boxes so a computer-use agent can emit a click or keystroke at them.Part 14
H
HallucinationNaming or describing something the pixels do not support, because the language prior outweighs the image evidence.Part 8
I
Image jailbreakAn optimised pixel perturbation that suppresses refusal behaviour a model learned from text.Part 16
Image tokensVision input expanded into context tokens; large, prefill-heavy and a cause of the prefill/decode inversion.Part 6
Inductive biasLocality and translation equivariance: a convolution has them, a patch transformer learns them from data.Part 1
InfoNCEThe symmetric contrastive loss on a similarity matrix that trains CLIP-style two-tower models.Part 2
K
KV cache (multimodal)The keys and values an image's tokens add; resident for the request and re-read on every decode step.Part 16
L
Latent world modelPredicts the next observation's representation so a policy can roll a plan forward in imagination.Part 15
M
Masked generation (MAR)Reveal positions in a seeded order and predict masked positions in parallel, with context on both sides.Part 13
Merge (2×2)Group a block of four patches into one token before the language model, dividing the token budget by four.Part 6
Modality embedderThe per-modality front end that maps a tokenizer's output into the trunk's width.Part 9
Modality-aware expertsA shared attention stack with a routed feed-forward pool that can specialise per modality.Part 9
N
Native dynamic resolutionPatch at the image's own size and aspect ratio, with 2-D position embeddings and a floating token count.Part 6
Next-scale (VAR)Predict whole scales coarse to fine, so sequential depth grows logarithmically rather than quadratically.Part 13
P
PatchifyCut an image into patches and flatten each into a vector of numbers the transformer can read.Part 1
Pixel-shuffleRepack a patch block into channels instead of averaging it, keeping detail at the same token count.Part 6
Position embeddingAdded to each patch so permutation-invariant attention knows where a patch came from; learned or sinusoidal.Part 1
PrefillRunning the language stack over image plus text tokens at once; compute-bound and carries the quadratic term.Part 16
Prompt injectionAn instruction a model reads from content rather than from the user; a screenshot is content.Part 16
Promptable segmentationOne trained model that returns a mask for a point, box or mask prompt without being fine-tuned per object.Part 14
ProjectorThe adapter that maps encoder vectors into the language model's embedding width; trained while the towers stay frozen.Part 5
Q
Q-FormerA querying bridge of a fixed set of learned queries between a frozen encoder and a language model.Part 5
Quantisation errorHalf a bin for a coordinate token; the precision floor no model can point finer than.Part 14
R
Raster orderRow-major autoregression over image tokens; nearest neighbours in space are far apart in sequence.Part 13
Referring expressionA phrase that selects one object by describing it and its relations to other objects.Part 14
S
Shared embedding spaceOne vector space where a picture and its caption land near each other, enabling retrieval and zero-shot tasks.Part 3
SigLIPA pairwise sigmoid contrastive objective, an alternative to the softmax InfoNCE loss.Part 2
Staged trainingThe freezing schedule and stage order — align, then fine-tune — that trains a vision-language model.Part 7
T
Temperature (contrastive)The scale that sharpens or flattens the softmax in contrastive training and in attention.Part 2
Token budgetThe number of image tokens a resolution scheme spends, and therefore its prefill and KV cost.Part 6
Token count$(H/P)^2$ for a square image; the sequence length the transformer actually sees.Part 1
U
Unified modelOne trunk that both reads images to text and generates images from text.Part 13
V
VLAA vision-language-action model: a multimodal trunk whose output head emits actions, discrete or continuous.Part 15
VQ-VAEAn autoencoder whose encoder vectors snap to a learned codebook of discrete entries.Part 13
Zero-shot classificationScoring an image against class-text embeddings in a shared space, with no classifier trained on those classes.Part 3
2

Architecture eras

From two towers to one omni trunk

The field's five stages are not replacements so much as a widening of what the interface between vision and language is allowed to be. Fusion here means where the two modalities first meet: in a thin bridge between two separate towers, in every attention layer of one shared trunk, or somewhere in between. Read the table by asking what each stage actually trains. The earlier the fusion, the more of the model is in the loop, meaning the more components receive gradients and can adapt to one another, and the more of the model has to be pretrained jointly. Early stages reuse strong pretrained halves and therefore train comparatively little; later stages pay for a joint run and buy direct cross-modal mixing in return.

StageRepresentative modelsWhat it trainsWhat it leaves unsolved
Two-tower contrastiveCLIP, SigLIPBoth towers, with a contrastive loss into one shared embedding spaceNo generation, no instruction following; a similarity score, not an answer
Connector / Q-FormerBLIP-2, Flamingo, InstructBLIPA small bridge over frozen encoders and a frozen language modelThe interface is the only thing that learns; the encoder never hears what the LLM needs
Adapter / projectorLLaVA, Qwen-VL, InternVLAn MLP projector, often with the language model, on a mostly frozen encoderResolution via tiling; the encoder stays fixed at whatever it was trained for
Early fusionChameleon, Gemini, FuyuOne trunk on interleaved text, image and audio tokens from the startNeeds a joint pretraining run; one shared, zero-sum token budget
OmniGPT-4o, Gemini, Qwen2.5-Omni, MoshiAny-to-any across text, image, audio and video, often with speech outputReal-time streaming and full-duplex turn-taking remain hard, and the token budget is shared by all
💡 Read the stage as an answer to "where does modality mixing happen". A bridge mixes in a thin interface; early fusion mixes inside every attention layer; an omni model mixes inside the trunk and across the output modalities as well. The difference is not cosmetic: a thin interface touches only the bridge, while mixing inside attention lets every modality influence every other token at every depth. Each step widens the loop and raises the pretraining bill.
3

Token budget cheat card

What each resolution scheme spends

An image is as many tokens as its patch grid says — the grid being the array of patch positions the encoder cut it into — and every scheme in the table is a different answer to the same question: keep the detail or keep the sequence short. The counts below use a 14-pixel patch, which is the standard CLIP/SigLIP geometry, and a 256-token per-tile block with a thumbnail for AnyRes, the thumbnail being a downsampled whole-image view that preserves global context while the tiles hold the detail. The arithmetic comes from Multimodal.tiles.nativeTokens, Multimodal.tiles.anyResTokens and Multimodal.cost.imageTokens; the FLOPs and bytes they imply — floating-point operations and memory traffic, the two quantities a serving bill is actually made of — are worked out in the resolution part.

SchemeInputTokensTrade
Fixed 224²224 × 224, P = 14256Constant cost, trivial batching; detail is lost in the resize
AnyRes 672²672 × 672, 336-px tiles12804 tiles + 1 thumbnail × 256; hard seams and a scaling cost
Native dynamic 448²448 × 448, P = 141024Aspect ratio kept, no seams; count floats with size
Native dynamic 896²896 × 896, P = 144096Doubling the side quadruples the tokens; the quadratic bites
2×2 merge on 448²448 × 448, P = 14, merge 2256Divide by four; pooling loses detail, pixel-shuffle keeps it in channels
⚠ The token count is billed twice. It sets the one-time prefill cost and the permanent KV cache the decode reads on every step, which is why compression at the vision side is worth more than the FLOPs alone suggest. Halving the tokens halves both bills. See the serving part for the full bill.
4

Evaluation benchmarks

What each measures, and how each leaks

The scores quoted in model cards — the short summary documents released alongside a model — come from a handful of benchmarks, where a benchmark is a fixed test set with a scoring rule. Each one has a specific way of being easier than the ability it names: some items can be guessed from the question alone, and some have leaked into training data. A high number with no contamination analysis is weaker evidence than a modest number with one, because the analysis is the only part that checks whether the test was genuinely unseen. Treat the middle column as the claim and the caveat column as the reason not to take it at face value.

BenchmarkWhat it measuresContamination and leakage caveats
MMMUCollege-level reasoning across diagrams, charts, tables and figuresKnowledge-heavy; many items are answerable from the question text and subject knowledge without the image
MathVistaMathematical reasoning grounded in visual contextsA fraction of items are solvable from text alone; contamination from public solution text is hard to rule out
DocVQAQuestion answering over scanned documentsMostly an OCR test; needs high resolution and is saturating on leading models
ChartQAReasoning about values and trends in chartsTemplate-generated charts and augmented training data create leakage between train and test
MMBenchMany abilities at once, scored by a language-model judgeJudge bias and answer-style sensitivity: the score moves with the judge as well as with the model
⚠ Direct hallucination measures are the cleaner test. CHAIR and POPE control the image and score the claim, so they isolate whether a model says things the pixels do not support; both build images with a known contents list, so a claim can be marked true or false against ground truth. The benchmark scores above are quoted more often and measure less cleanly; the hallucination part works through why.
📚

Further reading

Primary sources behind the numbers