Grounding
Describing a picture is easy to check: the words either match the scene or they do not. Pointing at one is harder, because a pointer is a pair of numbers and the model has to place those numbers exactly where the thing is. Grounding is the discipline of making a language model emit spatial coordinates — a box, a point, a mask — as tokens it already knows how to produce, so that "the red mug on the left" resolves to a rectangle a downstream system can act on. It is what turns a captioner into something that can be pointed at a robot, a document, or a button. The clearest way to hold the idea is as the difference between writing an answer and circling the answer on the page: a captioner writes, and a grounded model circles. Both say what is in the picture, but only the second says exactly where, and only a location is something another program can act on. That is what turns a captioner into something that can be pointed at a robot, a document, or a button. It is worth being precise about the word, because it is used loosely elsewhere: here grounding means tying a piece of language to a place in an image, nothing more mystical. This part follows the coordinate from a continuous number to a discrete token, then out to boxes, masks, referring expressions and a computer screen.
Coordinates as tokens
Clip a continuous pair onto a grid of ids
A language model emits tokens from a fixed vocabulary: a token is one symbol drawn from a finite list, and that list is the vocabulary, so everything the model can output is one of a known, countable set. A coordinate, by contrast, is a real number, a value on a continuous line with infinitely many possibilities between any two points. The standard reconciliation is to bin the image into a grid and let each coordinate be the index of a bin, where a bin is one cell of the grid and its index is the integer that names that cell. A box becomes four integers — two corners, each a row and a column — written as ordinary tokens like <bin_13>. A tiny example makes the snapping concrete: on a 32-by-32 grid, a corner a fraction 0.62 of the way across falls into column 19, and the model emits the token for that column rather than the exact value. The model then does what it is good at: it predicts the sequence of coordinate tokens that follows the referring phrase, exactly as it would predict the next word. The cost is resolution. A grid with $N$ bins per side places a corner within $1/N$ of its true position and no better, so an object smaller than a bin cannot be pointed at precisely, and a box's edges jitter by half a bin. The error is half a bin rather than a whole one because a token stands for the centre of its cell, so the worst a corner can be off is half the cell's width.
The number of bins is therefore a real budget decision, not a detail. More bins localise better but lengthen the vocabulary and give the model more ways to be wrong, because every extra column is another symbol for the softmax to spread probability over and another distinction the model has to learn from data. Fewer bins are easy to predict but coarser than the objects that matter, so a small mug and a distant bird can end up in the same cell. The trade cuts both ways, and neither end is free: precision costs vocabulary size, and vocabulary size costs statistical efficiency, since a model that must tell 256 columns apart needs more examples than one that must tell 8 apart. The demo quantises a point onto an $N \times N$ grid; the error arrow is the distance thrown away by the snap, and the readout names the coordinate token the model would emit.
A continuous point and its snapped cell centre on an $N \times N$ coordinate grid. The arrow is the quantisation error; each bin index is one coordinate token the model predicts.
Boxes and masks
A box is four tokens; a mask is many
Once a coordinate is a token, detection is a language task. Detection means finding objects in an image and returning a rectangle around each one, a bounding box, together with a label saying what it is; the name survives from the older computer-vision pipeline that did this with purpose-built heads, and the point here is that a language model can do it by emitting tokens instead. A grounded model is trained on text that interleaves an object's name with its normalised corner coordinates, so "mug" is followed by four bins and a score. Normalised means the coordinates are written as fractions of the image's width and height, running from 0 to 1, so that a token means the same relative place whatever the image's pixel size. Concretely, a training string might read the word mug, then the four bin indices for left, top, right and bottom, then a confidence value, and the loss is the ordinary next-token loss on that whole string. At inference — after training, when the model is answering rather than learning — it reads the phrase, emits the bins, and the bins are decoded back into a rectangle. Segmentation adds a third representation: rather than a box around the object, the model emits a polygon or a run of mask tokens that traces its outline. A polygon is a closed shape named by a short list of corner points, so it can follow a boundary that a rectangle cannot, while a mask is a dense per-pixel map saying, pixel by pixel, whether that pixel belongs to the object. The polygon is compact and the boundary is exact; the mask is dense and can represent a shape a box cannot, such as a ring or a curved handle. Both are still token sequences, which is the only property that matters to the trunk.
Confidence is the other half of the problem. A grounded model does not emit one box; it emits a ranked set, each box carrying a score, and a downstream system has to choose a threshold above which a box is trusted. That threshold is a single cut on the score: predictions above it are reported and predictions below it are thrown away. Set it low and the model reports everything that looks a little like a mug, so a fold of cloth or a patch of shadow becomes a false positive; set it high and it misses the one that is half-occluded, meaning partly hidden behind something else so only a fraction of it shows. The two failure modes pull in opposite directions, and the right cut depends on what each mistake costs: extra boxes are tolerable when a person reads the output, while they are not when a machine acts on it. The slider moves the threshold on a fixed set of predictions so the trade between recall and noise is visible, where recall is the fraction of the objects actually present that the model managed to report.
A procedural scene from GenMedia.img.circles with a fixed set of predicted boxes, thresholded by the slider. Colour runs from the least to the most confident prediction.
A high threshold keeps only the confident boxes and misses occluded objects; a low one keeps the recall and reports shapes that were never there. There is no threshold that is right for every image.
Referring expressions and promptable segmentation
A phrase in, a region out; a click in, a mask out
A referring expression is a phrase that picks out one object by describing it — "the plant in the corner", "the cup behind the book". The phrase does two jobs at once: it names a kind of thing, and it uses a spatial or relational word to choose one instance of that kind. Resolving one requires both language and space, because the model has to identify the candidate objects, use the relation to break the tie, and return the one box the phrase names. If a scene holds three mugs, the word "mug" alone does not single one out; "the mug on the left" does, and only if the model can compare positions. A referring benchmark is harder than a detection benchmark for exactly that reason, where a benchmark is a fixed test set with known right answers: the answer is not any mug, it is the mug the sentence chose, and a model that only detects objects has no way to respect the choice.
Promptable segmentation takes the complementary approach. Instead of a phrase, it accepts a prompt — a point inside an object, a box around it, or a rough mask — and returns a mask for whatever the prompt touches. A click is the simplest kind of prompt, an implicit statement that the point belongs to the thing being chosen. The prompt encoder turns the click into a vector, mapping raw coordinates and a little context into the same numeric form the image side already uses; a small mask decoder combines it with the image embeddings, the per-patch vectors the vision encoder produced, and the result is a region. The power is that one model answers many prompts without retraining: point at the handle and it segments the mug; box the page and it segments the text. The weights are identical across those two calls, which is what makes the interaction feel like one tool rather than a series of separate models. Click anywhere on the scene below and the nearest object is masked; the slider selects an object by its referring expression instead.
Click a point on the scene to segment the nearest object, or move the slider to resolve a referring expression. The soft region is the mask the prompt returns.
A click is an implicit positive prompt: it says "this point is inside the thing". A box adds a boundary, and a negative point says "not this". The same decoder handles all three.
Grounding a computer interface
The same coordinates, on a screen that happens to be an image
A screenshot is an image and every button on it has a rectangle, so grounding a computer interface is the grounding problem with a different vocabulary of objects: search fields, icons, list rows, submit buttons. The categories change, but the machinery does not, because a button is an object with a boundary like any other. A computer-use agent reads an instruction, grounds the element the instruction names to a box, and emits a click at that box's centre or a type action into it. The pixels are the observation and the coordinates are the action interface: the model never reaches into the underlying program, it only sees the rendered picture and it only acts by aiming at places in that picture. That is why vision-language-action models and GUI agents share their machinery — both turn a sentence into a target location and then into a movement, with coordinate tokens doing the work in between.
What changes is the precision the task demands. A mug box that is a few pixels off is still a correct caption; a submit-button box that is a few pixels off is a click on nothing, or worse on the wrong control. The tolerance for error shrinks because the consequence of error is an action rather than a description. Interface elements are also densely packed and text-heavy, so the model has to read labels as well as locate them: knowing that a rectangle holds a button is not enough if the model cannot tell "Submit" from "Cancel". On top of that, it must be robust to the same element appearing in a different place on the next screen, since layouts shift with window size, theme and content, and a model that memorised absolute positions would fail the moment the window changed. The schematic below marks the groundable elements of a page and highlights the one an instruction selects.
A schematic interface with its groundable elements boxed. The slider picks the element the instruction refers to; the arrow runs from the phrase to the grounded box.
The model must emit the box, not a pixel colour or a DOM id. Everything it knows about the page arrives through the screenshot, which is exactly the setting a grounded model was trained for.
Where this shows up
One representation, many interfaces
Detection and segmentation as language
Open-vocabulary detectors and promptable segmenters express their output as boxes, points and masks that a language model can emit token by token. Open-vocabulary means the set of things a model can find is not fixed in advance: because a class label is itself text, the detector can be asked for any object a phrase can name, including one it never saw as a labelled category. The unified models of the previous part provide the trunk; grounding is the task that teaches it where things are.
The interface for robots and agents
A grounded box is the bridge from a sentence to a movement. It is the intermediate step that makes the sentence usable by a controller, since "pick up the mug" is not something an arm can execute but a rectangle around the mug is. The next part takes the coordinate one step further, into action tokens a policy emits to move an arm or a cursor, where a policy is the component that maps what the model currently sees to an action.
Grounding also shows up wherever an interface is being automated — document parsing that has to return a field's bounding box, meaning software that reads a form and reports the rectangle around each entry; accessibility tooling that labels controls, so a screen reader can announce what a control is and where it sits; and evaluation of multimodal models on whether they can point rather than merely describe. A model that names the right object in prose but marks the wrong rectangle has failed a different test, and a stricter one. The resolution and merge choices of the token-budget part set the ceiling on all of it: a coordinate cannot be more precise than the grid the image was patched onto, so the finest box a model can ever emit is bounded by a decision made long before grounding is trained.
Further reading
These papers introduce the segmentation model that made prompting standard, the multimodal models that emit coordinates as text, and the studies of grounding an interface as an action space. Read together they trace the same idea from three directions: a promptable segmenter showing how much one model can do without retraining, grounded language models showing coordinates becoming ordinary tokens, and interface agents showing what changes when a coordinate becomes a click.
- Alexander Kirillov, Eric Mintun, Nikhila Ravi and coauthors, "Segment Anything", 2023 — the promptable segmentation model, its prompt encoder and its mask decoder, and the point and box prompts the demo imitates.
- Zhiliang Peng, Wenhui Wang, Li Dong and coauthors, "Kosmos-2: Grounding Multimodal Large Language Models to the World", 2023 — spatial tokens and the grounded box notation this part quantises.
- Keqin Chen, Zhao Zhang, Weili Zeng and coauthors, "Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic", 2023 — referring expressions and coordinate output from a chat model.
- Jianwei Yang, Hao Zhang, Feng Li and coauthors, "Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V", 2023 — overlaying numbered regions so a model can refer to them by id.
- Boyuan Zheng, Boyu Gou, Jihyung Kil and coauthors, "GPT-4V(ision) is a Generalist Web Agent, if Grounded", 2024 — grounding interface elements as boxes and acting on them, and the failure when the box is wrong.
Cheat sheet
| Term | Meaning here |
|---|---|
| Coordinate token | A quantised bin index standing in for a continuous coordinate |
| Bin resolution | Bins per side; more localises better but grows the vocabulary and the ways to be wrong |
| Grounded box | Four coordinate tokens marking two corners of an object, emitted after its phrase |
| Mask / polygon | An outline or dense region that a box cannot express; still a token sequence |
| Confidence threshold | The cut above which a predicted box is trusted; trades recall against noise |
| Referring expression | A phrase that selects one object by describing it and its relations |
| Promptable segmentation | One model that returns a mask for a point, box or mask prompt without retraining |
| GUI grounding | Localising interface elements to boxes so a click or keystroke can be emitted |
| Action interface | Coordinates as the boundary between perception and a real action on a screen or robot |