Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

Coordinates as tokens

Clip a continuous pair onto a grid of ids

A language model emits tokens from a fixed vocabulary: a token is one symbol drawn from a finite list, and that list is the vocabulary, so everything the model can output is one of a known, countable set. A coordinate, by contrast, is a real number, a value on a continuous line with infinitely many possibilities between any two points. The standard reconciliation is to bin the image into a grid and let each coordinate be the index of a bin, where a bin is one cell of the grid and its index is the integer that names that cell. A box becomes four integers — two corners, each a row and a column — written as ordinary tokens like <bin_13>. A tiny example makes the snapping concrete: on a 32-by-32 grid, a corner a fraction 0.62 of the way across falls into column 19, and the model emits the token for that column rather than the exact value. The model then does what it is good at: it predicts the sequence of coordinate tokens that follows the referring phrase, exactly as it would predict the next word. The cost is resolution. A grid with $N$ bins per side places a corner within $1/N$ of its true position and no better, so an object smaller than a bin cannot be pointed at precisely, and a box's edges jitter by half a bin. The error is half a bin rather than a whole one because a token stands for the centre of its cell, so the worst a corner can be off is half the cell's width.

The number of bins is therefore a real budget decision, not a detail. More bins localise better but lengthen the vocabulary and give the model more ways to be wrong, because every extra column is another symbol for the softmax to spread probability over and another distinction the model has to learn from data. Fewer bins are easy to predict but coarser than the objects that matter, so a small mug and a distant bird can end up in the same cell. The trade cuts both ways, and neither end is free: precision costs vocabulary size, and vocabulary size costs statistical efficiency, since a model that must tell 256 columns apart needs more examples than one that must tell 8 apart. The demo quantises a point onto an $N \times N$ grid; the error arrow is the distance thrown away by the snap, and the readout names the coordinate token the model would emit.

A continuous point and its snapped cell centre on an $N \times N$ coordinate grid. The arrow is the quantisation error; each bin index is one coordinate token the model predicts.

💡 By the end of this part you'll be able to explain how a box becomes four coordinate tokens, describe what a promptable segmentation model does with a point or a box (promptable meaning one trained model answers many prompts without being retrained, and segmentation meaning deciding which pixels belong to an object), quantise a referring expression into a box, and say why grounding a GUI — a graphical user interface, the buttons and fields on a screen — is the same problem with buttons instead of mugs.
2

Boxes and masks

A box is four tokens; a mask is many

Once a coordinate is a token, detection is a language task. Detection means finding objects in an image and returning a rectangle around each one, a bounding box, together with a label saying what it is; the name survives from the older computer-vision pipeline that did this with purpose-built heads, and the point here is that a language model can do it by emitting tokens instead. A grounded model is trained on text that interleaves an object's name with its normalised corner coordinates, so "mug" is followed by four bins and a score. Normalised means the coordinates are written as fractions of the image's width and height, running from 0 to 1, so that a token means the same relative place whatever the image's pixel size. Concretely, a training string might read the word mug, then the four bin indices for left, top, right and bottom, then a confidence value, and the loss is the ordinary next-token loss on that whole string. At inference — after training, when the model is answering rather than learning — it reads the phrase, emits the bins, and the bins are decoded back into a rectangle. Segmentation adds a third representation: rather than a box around the object, the model emits a polygon or a run of mask tokens that traces its outline. A polygon is a closed shape named by a short list of corner points, so it can follow a boundary that a rectangle cannot, while a mask is a dense per-pixel map saying, pixel by pixel, whether that pixel belongs to the object. The polygon is compact and the boundary is exact; the mask is dense and can represent a shape a box cannot, such as a ring or a curved handle. Both are still token sequences, which is the only property that matters to the trunk.

Confidence is the other half of the problem. A grounded model does not emit one box; it emits a ranked set, each box carrying a score, and a downstream system has to choose a threshold above which a box is trusted. That threshold is a single cut on the score: predictions above it are reported and predictions below it are thrown away. Set it low and the model reports everything that looks a little like a mug, so a fold of cloth or a patch of shadow becomes a false positive; set it high and it misses the one that is half-occluded, meaning partly hidden behind something else so only a fraction of it shows. The two failure modes pull in opposite directions, and the right cut depends on what each mistake costs: extra boxes are tolerable when a person reads the output, while they are not when a machine acts on it. The slider moves the threshold on a fixed set of predictions so the trade between recall and noise is visible, where recall is the fraction of the objects actually present that the model managed to report.

A procedural scene from GenMedia.img.circles with a fixed set of predicted boxes, thresholded by the slider. Colour runs from the least to the most confident prediction.

A high threshold keeps only the confident boxes and misses occluded objects; a low one keeps the recall and reports shapes that were never there. There is no threshold that is right for every image.

⚠ A box is a lossy answer. Four coordinates cannot say that the mug is behind the book, or that only its handle is visible; two different shapes can share one rectangle, and one rectangle can hide several objects. That is why segmentation exists, and why a box prompt and a mask output are often used together: the box localises crudely and the mask refines what the box enclosed.
3

Referring expressions and promptable segmentation

A phrase in, a region out; a click in, a mask out

A referring expression is a phrase that picks out one object by describing it — "the plant in the corner", "the cup behind the book". The phrase does two jobs at once: it names a kind of thing, and it uses a spatial or relational word to choose one instance of that kind. Resolving one requires both language and space, because the model has to identify the candidate objects, use the relation to break the tie, and return the one box the phrase names. If a scene holds three mugs, the word "mug" alone does not single one out; "the mug on the left" does, and only if the model can compare positions. A referring benchmark is harder than a detection benchmark for exactly that reason, where a benchmark is a fixed test set with known right answers: the answer is not any mug, it is the mug the sentence chose, and a model that only detects objects has no way to respect the choice.

Promptable segmentation takes the complementary approach. Instead of a phrase, it accepts a prompt — a point inside an object, a box around it, or a rough mask — and returns a mask for whatever the prompt touches. A click is the simplest kind of prompt, an implicit statement that the point belongs to the thing being chosen. The prompt encoder turns the click into a vector, mapping raw coordinates and a little context into the same numeric form the image side already uses; a small mask decoder combines it with the image embeddings, the per-patch vectors the vision encoder produced, and the result is a region. The power is that one model answers many prompts without retraining: point at the handle and it segments the mug; box the page and it segments the text. The weights are identical across those two calls, which is what makes the interaction feel like one tool rather than a series of separate models. Click anywhere on the scene below and the nearest object is masked; the slider selects an object by its referring expression instead.

Click a point on the scene to segment the nearest object, or move the slider to resolve a referring expression. The soft region is the mask the prompt returns.

A click is an implicit positive prompt: it says "this point is inside the thing". A box adds a boundary, and a negative point says "not this". The same decoder handles all three.

💡 Promptable means amortised. Training one model that accepts points, boxes and masks is what makes interactive segmentation usable: the user refines the region with successive clicks and the model never has to be fine-tuned for the particular object. Amortised here means the cost is paid once, at training time, rather than again for every object, and fine-tuning would mean updating the weights each time, which is the expense this design removes.
4

Grounding a computer interface

The same coordinates, on a screen that happens to be an image

A screenshot is an image and every button on it has a rectangle, so grounding a computer interface is the grounding problem with a different vocabulary of objects: search fields, icons, list rows, submit buttons. The categories change, but the machinery does not, because a button is an object with a boundary like any other. A computer-use agent reads an instruction, grounds the element the instruction names to a box, and emits a click at that box's centre or a type action into it. The pixels are the observation and the coordinates are the action interface: the model never reaches into the underlying program, it only sees the rendered picture and it only acts by aiming at places in that picture. That is why vision-language-action models and GUI agents share their machinery — both turn a sentence into a target location and then into a movement, with coordinate tokens doing the work in between.

What changes is the precision the task demands. A mug box that is a few pixels off is still a correct caption; a submit-button box that is a few pixels off is a click on nothing, or worse on the wrong control. The tolerance for error shrinks because the consequence of error is an action rather than a description. Interface elements are also densely packed and text-heavy, so the model has to read labels as well as locate them: knowing that a rectangle holds a button is not enough if the model cannot tell "Submit" from "Cancel". On top of that, it must be robust to the same element appearing in a different place on the next screen, since layouts shift with window size, theme and content, and a model that memorised absolute positions would fail the moment the window changed. The schematic below marks the groundable elements of a page and highlights the one an instruction selects.

A schematic interface with its groundable elements boxed. The slider picks the element the instruction refers to; the arrow runs from the phrase to the grounded box.

The model must emit the box, not a pixel colour or a DOM id. Everything it knows about the page arrives through the screenshot, which is exactly the setting a grounded model was trained for.

⚠ A grounded click is a real action. Once a model can turn a phrase into a coordinate and a coordinate into a click, the output is no longer text that a person can read and reject. That is what makes the safety discussion in the serving part of this volume concrete: an injected instruction in a screenshot can be grounded and executed like any other.
5

Where this shows up

One representation, many interfaces

Perception

Detection and segmentation as language

Open-vocabulary detectors and promptable segmenters express their output as boxes, points and masks that a language model can emit token by token. Open-vocabulary means the set of things a model can find is not fixed in advance: because a class label is itself text, the detector can be asked for any object a phrase can name, including one it never saw as a labelled category. The unified models of the previous part provide the trunk; grounding is the task that teaches it where things are.

Action

The interface for robots and agents

A grounded box is the bridge from a sentence to a movement. It is the intermediate step that makes the sentence usable by a controller, since "pick up the mug" is not something an arm can execute but a rectangle around the mug is. The next part takes the coordinate one step further, into action tokens a policy emits to move an arm or a cursor, where a policy is the component that maps what the model currently sees to an action.

Grounding also shows up wherever an interface is being automated — document parsing that has to return a field's bounding box, meaning software that reads a form and reports the rectangle around each entry; accessibility tooling that labels controls, so a screen reader can announce what a control is and where it sits; and evaluation of multimodal models on whether they can point rather than merely describe. A model that names the right object in prose but marks the wrong rectangle has failed a different test, and a stricter one. The resolution and merge choices of the token-budget part set the ceiling on all of it: a coordinate cannot be more precise than the grid the image was patched onto, so the finest box a model can ever emit is bounded by a decision made long before grounding is trained.

Further reading

These papers introduce the segmentation model that made prompting standard, the multimodal models that emit coordinates as text, and the studies of grounding an interface as an action space. Read together they trace the same idea from three directions: a promptable segmenter showing how much one model can do without retraining, grounded language models showing coordinates becoming ordinary tokens, and interface agents showing what changes when a coordinate becomes a click.

Cheat sheet

TermMeaning here
Coordinate tokenA quantised bin index standing in for a continuous coordinate
Bin resolutionBins per side; more localises better but grows the vocabulary and the ways to be wrong
Grounded boxFour coordinate tokens marking two corners of an object, emitted after its phrase
Mask / polygonAn outline or dense region that a box cannot express; still a token sequence
Confidence thresholdThe cut above which a predicted box is trusted; trades recall against noise
Referring expressionA phrase that selects one object by describing it and its relations
Promptable segmentationOne model that returns a mask for a point, box or mask prompt without retraining
GUI groundingLocalising interface elements to boxes so a click or keystroke can be emitted
Action interfaceCoordinates as the boundary between perception and a real action on a screen or robot
7

Check your understanding

0/4 answered