Why vision-language models hallucinate
A vision-language model is a system that takes an image as input and writes text about it, and if you ask one what is in a picture it will often describe a perfectly plausible scene that is not the one in front of it: a person who is not there, a cup that is only a shadow, a caption lifted from the distribution of captions rather than from the pixels. That failure has a name. A hallucination is a confident, fluent statement that the evidence does not support — here, text about an image that the image itself does not justify. This is not a decoding accident, meaning a one-off slip produced when the model chooses its next word, and it is not limited to a few bad checkpoints, meaning a few unlucky saved sets of trained weights. It follows from the way these models are built. The text half is a language model with a strong language prior — a learned expectation about which words and objects usually go together, distilled from an enormous amount of text — and that prior is conditioned on an image whose tokens it may only partly use. The analogy to keep in mind is autocomplete finishing a plausible sentence: given "a man sat down at the table and picked up his", the next word is almost certainly a fork or a glass, whether or not those things are in the scene you actually asked about. This part separates the two forces, the prior and the pixels; shows where attention actually goes; and looks at the benchmarks — standardized test sets used to score models — that try to measure the difference.
What hallucinations are
Confident, fluent, and not in the picture
An object hallucination is a named object that does not appear in the image, stated with the fluency of something that does. If the model says "a dog on the sofa" and there is no dog, that is an object hallucination even though the sentence is grammatical and confident. It is worth separating from a wrong answer, because the mechanism is different, and the fix may be different too. A wrong answer can come from a failure to understand the question, from a genuinely ambiguous image, or from a reasoning error; a clearer question or a better picture might repair it. A hallucination is specifically a claim that outruns the evidence in the pixels: the model is not confused about the answer, it is asserting content the image never supplied. Models produce them most readily for objects that are common in their training captions and plausible in the scene — a phone on a desk, a person at a table — because those are exactly the continuations the language model prefers. That is the autocomplete effect once more: the sentence is easy to finish, and ease of finishing is not the same thing as truth.
Three families are usually distinguished, and naming them helps you diagnose which one you are looking at. Object hallucination invents or misnames a thing, such as reporting a laptop where there is only a book. Attribute hallucination gets the colour or count wrong, such as calling three cups two, or describing a blue shirt as red. Relation hallucination puts real objects in the wrong spatial arrangement, such as saying the cup is to the left of the plate when it is to the right. Notice that in the last two cases every object named may genuinely be present; the error lies in a property of those objects or in the geometry between them, which makes these failures harder to catch by checking the object list alone. All three share a tell: the model's statement is more consistent with typical training text than with this particular image. The scene below makes the point with an abstract picture and a caption; drag the prior slider and watch the caption acquire an object the picture does not contain, while the picture itself never changes.
A procedural scene from GenMedia.img.circles drawn with GenMedia.raster. Green labels name objects that are present; red labels are objects the language prior supplies.
A low prior leaves the caption grounded in the blobs that are actually there; a high prior lets the model complete the scene with what it expects to see.
The language prior
What the model expects, against what the pixels say
A vision-language model's answer is not a function of the image alone. It is a function of two probability distributions: what the language model expects to be talking about, and what the image tokens provide as evidence. One way to picture this is to imagine the model holding two score sheets at once. The first is the language prior: how likely each object is to be the subject of the sentence, judged from text alone. The second is the pixel evidence: how strongly the image tokens support that object. To produce an answer the model has to combine them. When the two agree, the model is confident and usually right. When they disagree, the answer depends on how strongly each is weighted, and that weighting is not a design choice made by the practitioner — it is baked in by the pretraining, the data mixture and the fine-tuning. Pretraining is the earlier, large-scale training that gave the model its general knowledge; the data mixture is the blend of examples used in the later stages; fine-tuning is the final, smaller round of training on task data. The bars below make the disagreement explicit. The blue bars are the language prior over six candidate objects; the pink bars are the evidence in the image. The marker is the model's combined claim as the prior is turned up, so you can watch the winning object flip from one the pixels support to one only the text expects.
There is an asymmetry that makes the failure systematic rather than random. The prior is shared across every image and was trained on vastly more text than the visual pathway was trained on images, so it is the older and more practised of the two habits. The evidence is narrow and specific, and it is compressed twice before the language model ever sees it: once by the vision tower, the half of the system that turns pixels into vectors, and again by the merge, the step that condenses a large grid of image tokens into a smaller set. By the time the language model reads the image, much of the fine detail is already gone, and it is working from a short summary rather than the full picture. A prior that is merely a little stronger therefore wins on exactly the objects that are plausible but absent, which is the definition of a hallucination. Grounding training — teaching the model to point at where in the image its answer comes from, as if circling the answer on the page — and preference tuning, which trains on pairs of better and worse answers, are both attempts to rebalance this equation; neither removes it, because neither changes the fact that the text habit is the broader one.
Language prior and pixel evidence over the candidate objects, with the model's combined claim marked. The prior and the evidence are each normalised, then mixed geometrically by the slider.
language prior pixel evidence combined claim
Attention over image tokens
Where the model actually looks
If the language prior were the whole story, a model that attends carefully to the image would be immune. It is not, because attention itself is uneven. Attention is the mechanism by which each token decides how much to read from every other token, and the weights it assigns form a probability distribution: they are all positive and they sum to one, so attention is a fixed budget that has to be spent somewhere. Transformer attention has a well-documented tendency to park that budget on a few positions — usually the first token or two, and some special tokens — that carry little information. These attention sinks act as a place to put "no opinion": a spot the model can dump weight on when it has nothing specific to look at. The behaviour is not a flaw the designers intended but a consequence of the softmax, which has to distribute a row of weights across every available position even when none of them is a useful match; over training the model turns a few low-information positions into a convenient resting spot. The behaviour was first described in text-only models and it persists into multimodal ones, where the sequence begins with a long block of image tokens. A few image tokens soak up a large share of the attention, and the tokens that contain the object you asked about are left with a thin slice of the remaining mass. The model looks like it is attending to the image; in aggregate it is attending mostly to itself.
Because the weights in a row sum to one, a large weight in one column has to come out of the others, which is why a sink is not a harmless curiosity. The heatmap below shows attention from text queries to image tokens, with a few sink columns. Turn the sink strength up and the real columns fade; that is the mechanism by which a correct answer to an easy question coexists with an invented object in the same caption. The model can read the easy feature and still miss the small object, because the easy feature happened to sit in a column that survived the sinks. The practical consequence is that adding more image tokens does not automatically make the model use them — at some point the added tokens dilute the ones that matter, spreading the same limited budget across a longer sequence.
Attention from text queries (rows) to 64 image tokens (columns), from a seeded stream. Sharp yellow columns are attention sinks; the rest is the evidence the model is actually reading.
Sinks are a normal property of trained transformers, not a bug in one model. What varies is how much of the image's information survives them.
Measuring it
A benchmark measures the task, and the task leaks
Evaluation means scoring a model on a fixed task so that two models can be compared on the same footing, and hallucination is measured in two ways: directly, by asking a model to describe an image and counting the objects that are not there, and indirectly, through benchmarks whose questions cannot be answered without looking. The direct methods — CHAIR, POPE and their descendants — are the cleanest, because they control the image and score the claim. CHAIR, short for Caption Hallucination Assessment with Image Relevance, takes a caption the model wrote and divides the number of invented objects by the number of objects it mentioned. POPE, short for Polling-based Object Probing Evaluation, asks a long list of yes-or-no questions of the form "is there a chair?", deliberately mixing objects that are present, objects that are absent, and objects that merely tend to co-occur. The indirect methods are the ones quoted in model cards, and each of them has a specific way of leaking. A multiple-choice question with a plausible option can sometimes be answered from the text of the question alone; a chart question drawn from a template lets a language model guess from the title; a document benchmark can be partly solved by an OCR system — optical character recognition, which reads text out of an image — with no vision-language reasoning at all. The price of the direct methods is that they need a reference caption and a trusted list of which objects are truly present, so they measure caption fluency as well as honesty; the indirect methods avoid that labour, but as a result they leave room to be passed without looking at the picture. A benchmark that can be gamed this way still produces a number, which is exactly why the number needs a caveat attached.
| Benchmark | What it measures | Main caveat |
|---|---|---|
| MMMU | College-level reasoning across diagrams, charts, tables and figures | Knowledge-heavy; many items are answerable from the question text and subject knowledge without the image |
| MathVista | Mathematical reasoning grounded in visual contexts | Solvable from text for a fraction of items; contamination from public solution text is hard to rule out |
| DocVQA | Question answering over scanned documents | Mostly an OCR test; needs high resolution, and saturating on the leading models |
| ChartQA | Reasoning about values and trends in charts | Template-generated charts and augmented training data create leakage between train and test |
| MMBench | Many abilities, scored by a language-model judge | Judge bias and answer-style sensitivity; the score moves with the judge as well as the model |
Where this shows up
Every part of the pipeline contributes
The token budget sets the ceiling
The merge and tiling choices decide how much pixel evidence survives to the language model. Compress aggressively and the small object the question is about may be one vector among many, averaged together with everything around it; the prior then has less to argue against, because the signal that would have contradicted it has been smoothed away. This is the fixed-size suitcase problem from the resolution discussion: a tighter pack fits into fewer tokens, but whatever you leave out cannot be recovered later.
The mixture sets the prior
The data mixture is what fixes the language prior, and preference tuning is the rebalancing act that follows. A model trained mostly on short captions hallucinates differently from one trained on long instruction answers: the caption model tends to name the most typical object for the scene, while the instruction model tends to keep talking and to elaborate. That is why a hallucination is a clue about the training data, not only about the image in front of the model. Reasoning-tuned VLMs bought with group-relative RL — reinforcement learning that scores each attempt against the other samples drawn for the same prompt rather than against a fixed target — trade hallucination behaviour for accuracy and verbosity: they spend more tokens and more latency, and the extra reasoning both catches some invented objects and talks itself into others, so the hallucination rate is a property of the reward mix, not only the data mixture, as the RL stage makes plain.
The interface is implicated too: the bridge decides how much of the vision tower's signal reaches the language model at all. A projector passes the vectors through more or less intact, whereas a Q-Former — a small module that reads the image with a fixed set of learned queries — returns only as many vectors as it has queries, so its query count is a hard cap on how many distinct things the model can be told. When a hallucination appears, the honest ordering of suspects is the mixture, the token budget, the bridge, and only then the decoding settings, where decoding means the choices made at generation time, such as how the next word is sampled. That order matters because decoding settings change the packaging of an answer far more than they change what the model knows: sampling more cautiously can make a model repeat itself or hedge, but it cannot put an object into the image.
Further reading
The work below defines object hallucination, measures it with direct probes, traces it to the language prior, and shows how attention sinks behave in trained transformers. Read them in that order and the arc runs from a careful definition, to a benchmark that counts the errors, to a diagnosis of why they happen, and finally to the mechanism inside attention that lets some image tokens be ignored.
- Anna Rohrbach, Lisa Anne Hendricks, Kaylee Burns and coauthors, "Object Hallucination in Image Captioning", 2018 — CHAIR, the first systematic count of objects that were never in the image.
- Yifan Li, Yifan Du, Kun Zhou and coauthors, "Evaluating Object Hallusion in Large Vision-Language Models", 2023 — POPE, polling-based probing that separates presence, absence and co-occurrence.
- Zhiqing Xiao, Qingyan Guo, Yancheng He and coauthors, "Can We See Like a Language Model? A Survey on Hallucination in Large Vision-Language Models", 2024 — a taxonomy of object, attribute and relation hallucination and the training-side causes.
- Guangxuan Xiao, Yuandong Tian, Beidi Chen and coauthors, "Efficient Streaming Language Models with Attention Sinks", 2023 — the sink phenomenon that multimodal attention inherits.
- Xiang Yue, Yuansheng Ni, Kai Zhang and coauthors, "MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark", 2023 — the benchmark whose knowledge-heavy items make the text-prior caveat concrete.
Cheat sheet
| Term | Meaning here |
|---|---|
| Object hallucination | Naming a thing that is not in the image |
| Attribute hallucination | Getting a colour, material or count wrong |
| Relation hallucination | Real objects placed in the wrong spatial relationship |
| Language prior | The model's expectation of what is being discussed, from text pretraining and the data mixture |
| Pixel evidence | What the image tokens actually support, after compression and the bridge |
| Attention sink | A few positions that absorb a large share of attention without carrying information |
| CHAIR / POPE | Direct hallucination measures; POPE polls presence, absence and co-occurrence |
| Contamination | Benchmark items present in training data, so the score measures recall, not ability |
| Judge bias | A model-scored benchmark moves when the judge changes, not only the model |