Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

What hallucinations are

Confident, fluent, and not in the picture

An object hallucination is a named object that does not appear in the image, stated with the fluency of something that does. If the model says "a dog on the sofa" and there is no dog, that is an object hallucination even though the sentence is grammatical and confident. It is worth separating from a wrong answer, because the mechanism is different, and the fix may be different too. A wrong answer can come from a failure to understand the question, from a genuinely ambiguous image, or from a reasoning error; a clearer question or a better picture might repair it. A hallucination is specifically a claim that outruns the evidence in the pixels: the model is not confused about the answer, it is asserting content the image never supplied. Models produce them most readily for objects that are common in their training captions and plausible in the scene — a phone on a desk, a person at a table — because those are exactly the continuations the language model prefers. That is the autocomplete effect once more: the sentence is easy to finish, and ease of finishing is not the same thing as truth.

Three families are usually distinguished, and naming them helps you diagnose which one you are looking at. Object hallucination invents or misnames a thing, such as reporting a laptop where there is only a book. Attribute hallucination gets the colour or count wrong, such as calling three cups two, or describing a blue shirt as red. Relation hallucination puts real objects in the wrong spatial arrangement, such as saying the cup is to the left of the plate when it is to the right. Notice that in the last two cases every object named may genuinely be present; the error lies in a property of those objects or in the geometry between them, which makes these failures harder to catch by checking the object list alone. All three share a tell: the model's statement is more consistent with typical training text than with this particular image. The scene below makes the point with an abstract picture and a caption; drag the prior slider and watch the caption acquire an object the picture does not contain, while the picture itself never changes.

A procedural scene from GenMedia.img.circles drawn with GenMedia.raster. Green labels name objects that are present; red labels are objects the language prior supplies.

A low prior leaves the caption grounded in the blobs that are actually there; a high prior lets the model complete the scene with what it expects to see.

💡 By the end of this part you'll be able to define object, attribute and relation hallucination, explain how a language prior can override pixel evidence, read an attention map for the sink behaviour that starves image tokens, and judge what a benchmark like MMMU or DocVQA actually measures. (Pixel evidence means what the image tokens actually support, and an attention map is a heatmap of how strongly each text query looks at each image token.)
2

The language prior

What the model expects, against what the pixels say

A vision-language model's answer is not a function of the image alone. It is a function of two probability distributions: what the language model expects to be talking about, and what the image tokens provide as evidence. One way to picture this is to imagine the model holding two score sheets at once. The first is the language prior: how likely each object is to be the subject of the sentence, judged from text alone. The second is the pixel evidence: how strongly the image tokens support that object. To produce an answer the model has to combine them. When the two agree, the model is confident and usually right. When they disagree, the answer depends on how strongly each is weighted, and that weighting is not a design choice made by the practitioner — it is baked in by the pretraining, the data mixture and the fine-tuning. Pretraining is the earlier, large-scale training that gave the model its general knowledge; the data mixture is the blend of examples used in the later stages; fine-tuning is the final, smaller round of training on task data. The bars below make the disagreement explicit. The blue bars are the language prior over six candidate objects; the pink bars are the evidence in the image. The marker is the model's combined claim as the prior is turned up, so you can watch the winning object flip from one the pixels support to one only the text expects.

There is an asymmetry that makes the failure systematic rather than random. The prior is shared across every image and was trained on vastly more text than the visual pathway was trained on images, so it is the older and more practised of the two habits. The evidence is narrow and specific, and it is compressed twice before the language model ever sees it: once by the vision tower, the half of the system that turns pixels into vectors, and again by the merge, the step that condenses a large grid of image tokens into a smaller set. By the time the language model reads the image, much of the fine detail is already gone, and it is working from a short summary rather than the full picture. A prior that is merely a little stronger therefore wins on exactly the objects that are plausible but absent, which is the definition of a hallucination. Grounding training — teaching the model to point at where in the image its answer comes from, as if circling the answer on the page — and preference tuning, which trains on pairs of better and worse answers, are both attempts to rebalance this equation; neither removes it, because neither changes the fact that the text habit is the broader one.

Language prior and pixel evidence over the candidate objects, with the model's combined claim marked. The prior and the evidence are each normalised, then mixed geometrically by the slider.

language prior  pixel evidence  combined claim

⚠ The prior is strongest where the evidence is weakest. Objects that are common in captions are also the ones the model is most willing to assert without support, because it has seen them in so many similar scenes. That is why the hallucinated objects are rarely strange; they are the expected ones. When you are auditing a model, that tells you to probe the ordinary furniture of a scene — the phone, the chair, the laptop — and not only the exotic details.
3

Attention over image tokens

Where the model actually looks

If the language prior were the whole story, a model that attends carefully to the image would be immune. It is not, because attention itself is uneven. Attention is the mechanism by which each token decides how much to read from every other token, and the weights it assigns form a probability distribution: they are all positive and they sum to one, so attention is a fixed budget that has to be spent somewhere. Transformer attention has a well-documented tendency to park that budget on a few positions — usually the first token or two, and some special tokens — that carry little information. These attention sinks act as a place to put "no opinion": a spot the model can dump weight on when it has nothing specific to look at. The behaviour is not a flaw the designers intended but a consequence of the softmax, which has to distribute a row of weights across every available position even when none of them is a useful match; over training the model turns a few low-information positions into a convenient resting spot. The behaviour was first described in text-only models and it persists into multimodal ones, where the sequence begins with a long block of image tokens. A few image tokens soak up a large share of the attention, and the tokens that contain the object you asked about are left with a thin slice of the remaining mass. The model looks like it is attending to the image; in aggregate it is attending mostly to itself.

Because the weights in a row sum to one, a large weight in one column has to come out of the others, which is why a sink is not a harmless curiosity. The heatmap below shows attention from text queries to image tokens, with a few sink columns. Turn the sink strength up and the real columns fade; that is the mechanism by which a correct answer to an easy question coexists with an invented object in the same caption. The model can read the easy feature and still miss the small object, because the easy feature happened to sit in a column that survived the sinks. The practical consequence is that adding more image tokens does not automatically make the model use them — at some point the added tokens dilute the ones that matter, spreading the same limited budget across a longer sequence.

Attention from text queries (rows) to 64 image tokens (columns), from a seeded stream. Sharp yellow columns are attention sinks; the rest is the evidence the model is actually reading.

Sinks are a normal property of trained transformers, not a bug in one model. What varies is how much of the image's information survives them.

💡 Attention maps explain, they do not excuse. A sparse or sink-heavy map tells you why an answer ignored a region, which is useful for debugging. It does not by itself prove the model used the tokens it attended to, and attention weight is not the same as causal influence: a weight records how much a query looked at a key, not whether the information behind that key changed the final answer.
4

Measuring it

A benchmark measures the task, and the task leaks

Evaluation means scoring a model on a fixed task so that two models can be compared on the same footing, and hallucination is measured in two ways: directly, by asking a model to describe an image and counting the objects that are not there, and indirectly, through benchmarks whose questions cannot be answered without looking. The direct methods — CHAIR, POPE and their descendants — are the cleanest, because they control the image and score the claim. CHAIR, short for Caption Hallucination Assessment with Image Relevance, takes a caption the model wrote and divides the number of invented objects by the number of objects it mentioned. POPE, short for Polling-based Object Probing Evaluation, asks a long list of yes-or-no questions of the form "is there a chair?", deliberately mixing objects that are present, objects that are absent, and objects that merely tend to co-occur. The indirect methods are the ones quoted in model cards, and each of them has a specific way of leaking. A multiple-choice question with a plausible option can sometimes be answered from the text of the question alone; a chart question drawn from a template lets a language model guess from the title; a document benchmark can be partly solved by an OCR system — optical character recognition, which reads text out of an image — with no vision-language reasoning at all. The price of the direct methods is that they need a reference caption and a trusted list of which objects are truly present, so they measure caption fluency as well as honesty; the indirect methods avoid that labour, but as a result they leave room to be passed without looking at the picture. A benchmark that can be gamed this way still produces a number, which is exactly why the number needs a caveat attached.

BenchmarkWhat it measuresMain caveat
MMMUCollege-level reasoning across diagrams, charts, tables and figuresKnowledge-heavy; many items are answerable from the question text and subject knowledge without the image
MathVistaMathematical reasoning grounded in visual contextsSolvable from text for a fraction of items; contamination from public solution text is hard to rule out
DocVQAQuestion answering over scanned documentsMostly an OCR test; needs high resolution, and saturating on the leading models
ChartQAReasoning about values and trends in chartsTemplate-generated charts and augmented training data create leakage between train and test
MMBenchMany abilities, scored by a language-model judgeJudge bias and answer-style sensitivity; the score moves with the judge as well as the model
⚠ Contamination changes the meaning of a score. If a benchmark's questions and answers were in the training data, the score measures recall of the benchmark rather than the ability it names. Web-scraped training corpora make this the default assumption until a provider shows a held-out or freshly collected set, where held-out means the test items were deliberately kept out of training so the model has never seen them. A high score with no contamination analysis is weaker evidence than a modest score with one. A modest score from a clean, freshly collected set is therefore real evidence about ability, while a large score from a contaminated set may only show that the test was seen before.
5

Where this shows up

Every part of the pipeline contributes

Compression

The token budget sets the ceiling

The merge and tiling choices decide how much pixel evidence survives to the language model. Compress aggressively and the small object the question is about may be one vector among many, averaged together with everything around it; the prior then has less to argue against, because the signal that would have contradicted it has been smoothed away. This is the fixed-size suitcase problem from the resolution discussion: a tighter pack fits into fewer tokens, but whatever you leave out cannot be recovered later.

Training

The mixture sets the prior

The data mixture is what fixes the language prior, and preference tuning is the rebalancing act that follows. A model trained mostly on short captions hallucinates differently from one trained on long instruction answers: the caption model tends to name the most typical object for the scene, while the instruction model tends to keep talking and to elaborate. That is why a hallucination is a clue about the training data, not only about the image in front of the model. Reasoning-tuned VLMs bought with group-relative RL — reinforcement learning that scores each attempt against the other samples drawn for the same prompt rather than against a fixed target — trade hallucination behaviour for accuracy and verbosity: they spend more tokens and more latency, and the extra reasoning both catches some invented objects and talks itself into others, so the hallucination rate is a property of the reward mix, not only the data mixture, as the RL stage makes plain.

The interface is implicated too: the bridge decides how much of the vision tower's signal reaches the language model at all. A projector passes the vectors through more or less intact, whereas a Q-Former — a small module that reads the image with a fixed set of learned queries — returns only as many vectors as it has queries, so its query count is a hard cap on how many distinct things the model can be told. When a hallucination appears, the honest ordering of suspects is the mixture, the token budget, the bridge, and only then the decoding settings, where decoding means the choices made at generation time, such as how the next word is sampled. That order matters because decoding settings change the packaging of an answer far more than they change what the model knows: sampling more cautiously can make a model repeat itself or hedge, but it cannot put an object into the image.

Further reading

The work below defines object hallucination, measures it with direct probes, traces it to the language prior, and shows how attention sinks behave in trained transformers. Read them in that order and the arc runs from a careful definition, to a benchmark that counts the errors, to a diagnosis of why they happen, and finally to the mechanism inside attention that lets some image tokens be ignored.

Cheat sheet

TermMeaning here
Object hallucinationNaming a thing that is not in the image
Attribute hallucinationGetting a colour, material or count wrong
Relation hallucinationReal objects placed in the wrong spatial relationship
Language priorThe model's expectation of what is being discussed, from text pretraining and the data mixture
Pixel evidenceWhat the image tokens actually support, after compression and the bridge
Attention sinkA few positions that absorb a large share of attention without carrying information
CHAIR / POPEDirect hallucination measures; POPE polls presence, absence and co-occurrence
ContaminationBenchmark items present in training data, so the score measures recall, not ability
Judge biasA model-scored benchmark moves when the judge changes, not only the model
7

Check your understanding

0/4 answered