Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

From seeing to acting

A language model with a different output head

The shortest description of a vision-language-action model is this: it is the model from the earlier parts of this volume with the language head replaced, or augmented, by one that emits actions. The language head is the final layer that turns the trunk's internal representation into a distribution over the next word, and here it is either retrained to emit numbers or joined by a new head that does. The image tower patches the camera view, cutting it into tiles and turning each tile into a vector, the trunk reads those vectors together with the instruction, and instead of predicting the next word it predicts the next command. Everything the trunk learned about objects, spatial relations and instructions transfers, because recognising that the red block is left of the blue one, and that "pick up" names an object, is exactly what a manipulation instruction needs. That is why starting from a pretrained checkpoint matters so much: teaching a vision-language trunk the world from scratch is expensive, while fine-tuning an existing one on action data costs a fraction of it and keeps the knowledge. The new part is only the mapping from the trunk's representation to a control signal, meaning the numbers a controller turns into motor movement.

The output representation is where the design splits. One family keeps the language interface exactly as it is and discretises the action, meaning each continuous number is rounded onto a fixed grid of bins and represented by the bin's index, so an action is a short sequence of tokens the trunk can emit with the same cross-entropy loss it uses for words. Cross-entropy is the standard training objective that rewards putting probability on the correct next symbol, and it applies only when the target is one of a finite set, which the bins provide. The other family leaves the action continuous and adds a small expert that predicts it with a regression objective, meaning the expert outputs real numbers and is trained by shrinking the distance between its output and the target, usually a flow-matching velocity. The two routes are not merely stylistic: they make different assumptions about whether an action is naturally a category that can be named or a quantity that has to be measured. The diagram shows the shared pipeline; the two following steps take the two output routes in turn.

The shared pipeline: camera and instruction into the vision-language trunk, then an action output that is either discretised into tokens or produced by a continuous action expert.

The trunk is unchanged from the rest of the volume. A VLA is a multimodal model whose output head is a controller, which is why it inherits both the strengths and the failure modes of the underlying VLM.

💡 By the end of this part you'll be able to explain how an action becomes a sequence of discretised tokens, describe what a flow-matching action expert predicts instead, say why a policy is asked for a chunk of actions rather than one, and name what a latent world model adds to the loop. A latent world model predicts the next observation in the trunk's internal vector form rather than as pixels, so a rollout of the future stays cheap.
2

Actions as tokens

Discretise a control signal and speak it

The first family borrows a trick from the tokenizer discussion: clip each dimension of a continuous action onto a fixed set of bins and emit the bin index as a token. Degrees of freedom are the independent axes a robot can move, so a seven-degree-of-freedom arm has seven of them, each driven by one number, and one action has seven dimensions. One action step is therefore seven tokens, one per axis, and a policy that wants to move for half a second emits a few dozen of them in a row, because commands are issued many times a second. The appeal is enormous — no new head, no new loss, no architectural change at all. The same trunk that can write "pick up the red block" can, after fine-tuning on action data, write the numbers that make it happen, where fine-tuning means continuing training on a smaller specialised set of robot demonstrations. It can also interleave reasoning text with the action tokens in one stream, so it can narrate what it is about to do and then do it.

The cost is quantisation, the rounding that replaces each exact value with the nearest bin's index and discards the difference. A bin represents a range of joint angles, so every emitted token carries an error of up to half a bin, and a controller that receives a coarse target has to smooth it or the motion stutters. The error is half a bin rather than a whole one because the token stands for the centre of its bin, so the worst a value can be off is half the bin's width. More bins reduce the error but lengthen the sequence, because every extra bin is another symbol the sampler can choose, and give the sampler more ways to emit an implausible combination; fewer bins are easy to predict but too coarse to be safe. Suppose a joint's range is divided into only sixteen bins: a command can then be off by a thirty-second of the joint's travel on either side, which at the end of a long reach becomes a visible wobble. The demo discretises a seven-dimensional action over an adjustable number of bins so the tension is visible lane by lane.

Seven action dimensions, each discretised into the number of bins on the slider. The highlighted cell per lane is the token the policy would emit; the trajectory step slider moves along a demonstration.

⚠ Quantisation is a safety budget, not just an accuracy one. A coarse bin is a real distance the arm can move in one command. The choice of bin count is therefore a choice about the worst-case error the controller must absorb, and it is set by the tokenizer before any policy is trained. Halving the bin count doubles the largest jump one command can ask for, so the smoothness of every motion is capped at tokenizer time.
3

The flow-matching expert

A continuous head that predicts velocity, not a category

The second family refuses to discretise at all. It keeps the action continuous and attaches a small action expert to the pretrained trunk: a compact network that reads the trunk's representation of the scene and instruction and produces a velocity field over action space. A velocity field is a rule that gives a direction and a speed at any point, action space is the space of possible actions with one coordinate per joint, and the expert's job is to say which way and how fast to move through it. To feel the idea before its name, picture sliding a cart along straight rails at constant speed. Suppose the target action sets a gripper width of 0.4, and a noise sample starts it at 0.9. The simplest possible route from 0.9 to 0.4 is a straight line travelled at a constant speed, so a single number — the velocity — describes the entire journey, and if the expert predicts it correctly the sample lands exactly on the target. Training uses flow matching — a noise sample and a target action are connected by a straight line, and the expert is trained to predict the constant velocity that carries one to the other. Sampling starts at noise and integrates that velocity for a handful of steps, which is why these policies can run at control rates that a hundreds-step diffusion sampler cannot.

The reason the straight path matters is the number of function evaluations, meaning how many times the expert network has to be called to produce one action. A curved path needs many small steps to follow accurately, because each step approximates the curve only over a short distance; a path the model has been trained to make straight can be traversed in a few Euler steps, where an Euler step is the simplest way to advance along a direction. Because the path is straight and the velocity is constant along it, the expert only has to be right about one vector rather than about a schedule of changing directions, and there is no compounding of per-step error the way a curve invites. A few evaluations per action chunk is the difference between a policy that fits a control loop and one that does not, where the control loop is the observe-decide-act cycle the controller repeats many times a second. The plot below draws the oracle straight paths from a cloud of noise samples to one target action, and the slider walks along the path so you can read the velocity, which is constant by construction.

Straight flow paths from Diffusion.flowX carrying noise samples to a target action, with the constant velocity of Diffusion.flowVelocity. The slider walks the flow time from noise to action.

A straight path is a property the training objective encourages, not an assumption: regressing x1 − x0 along a line is what lets a few integration steps reach the action.

💡 Why an expert rather than the trunk itself? A diffusion or flow head has its own parameters and its own loss, so it can be trained to be smooth without disturbing the discrete next-token objective the trunk was pretrained on. That separation matters because the trunk's objective wants sharp, confident predictions while a smooth continuous action needs the opposite. It is the unified-model settlement of the previous-to-last part, applied to control.
4

Chunking and world models

Move before the next thought finishes

A policy that emits one action per inference pays the full latency of a forward pass between every command. Inference is one run of the network, and latency is the wall-clock time that run takes. If the trunk takes fifty milliseconds and the controller runs at a hundred hertz, meaning a hundred commands per second, then the arm needs a fresh target every ten milliseconds while the model can supply one only every fifty, so the arm is idle most of the time. Action chunking fixes that by asking the model for a whole horizon of actions at once — the next sixteen or fifty steps — and executing them open-loop while the next inference runs in parallel. The horizon is the number of future steps requested, and open-loop means the arm keeps running the commands it was given without checking what actually happened. Chunking also smooths the motion, because a single inference fixes a consistent run of commands instead of re-deciding at every tick. The cost is reactivity: while the chunk is executed the policy cannot see a new observation, so a chunk that is too long responds slowly to a moved object and a chunk that is too short pays the inference tax on every step. Temporal ensembling, which blends overlapping chunks by averaging the predictions two successive inferences make for the same moment, is the usual compromise.

A latent world model addresses the other half of the loop. A world model is a network that predicts what happens next, and latent means it makes that prediction in the trunk's internal vector form rather than as a picture, which keeps the prediction cheap. If the trunk can predict the representation of the next observation as well as the action, then a policy can plan a short rollout in latent space, score candidate chunks against the predicted future, and act on the best one without touching the real environment. A rollout is that imagined sequence of steps, and scoring means comparing candidate actions against the future the model just predicted. The same prediction, trained on the same interleaved stream, tells the model what its action is likely to cause, which is what separates "emit a plausible movement" from "know what the movement will do". Planning against that predicted future is what lets a policy improve a chunk before it commits to it, which matters because a real arm cannot undo a movement once it is underway. The timeline below shows chunks being inferred, executed and re-planned as the horizon changes.

A control timeline: solid blocks are actions executed open-loop, the dashed extension is the part of the chunk a latent world model rolls forward, and the ticks are inference calls. The slider sets the chunk horizon in steps.

Longer chunks mean fewer inference calls and smoother motion, but slower reaction. The world model is what lets the policy spend the idle time imagining the future instead of waiting.

⚠ Open-loop execution is a commitment. Once a chunk is being executed the observation no longer steers it. Every design that chunks is choosing how much of the future to guess, and a world model is only useful if its predicted future is accurate enough to be worth planning against. The longer the chunk, the more of that guess is already spent by the time the world changes, so a bad prediction is expensive to correct.
5

Where this shows up

From a sentence to a movement

Discrete

Actions as language tokens

RT-2 discretises each action dimension and appends the tokens to the vocabulary, the model's fixed list of output symbols, so a web-scale vision-language model can be fine-tuned into a policy that still answers questions. A web-scale model is one pretrained on a very large and general mixture of text and images. The inheritance is the point: semantic understanding transfers with the tokens, so the robot carries a vocabulary of objects and relations it absorbed from ordinary text and images rather than from robot data alone.

Continuous

A flow-matching action expert

The pi-zero family keeps a pretrained vision-language backbone and adds a flow-matching action expert, trained to produce smooth, high-frequency action chunks — commands issued many times a second — that the backbone alone could not emit, because a trunk trained to pick a word does not readily output a continuous number. It is the continuous route of step three, paired with the chunking of step four.

The full loop is where the rest of the guide meets control: the same latent world models that generate video (the sibling volume) are the ones a policy rolls forward, and the same token-budget arithmetic decides whether the camera view and the instruction fit beside the action horizon in one context. Token budget is the number of tokens a fixed context window can hold, and the camera view competes with the instruction for that room. A VLA sits on top of a control stack — the layered software between a command and a motor.

Further reading

These papers are the two output routes — discretised action tokens and a continuous action expert — together with the world-model idea the loop borrows from. Read together they trace one action from a word-like token, through a continuous velocity, to a chunk planned against an imagined future.

Cheat sheet

TermMeaning here
VLAA multimodal trunk whose output head emits actions instead of, or beside, text
Discretised actionEach action dimension clipped onto bins, emitted as an ordinary token
Action tokenOne bin index; a seven-axis arm emits seven per action step
Bin countSets the worst-case command error and the length of the action sequence
Action expertA small continuous head on the trunk, trained with regression rather than cross-entropy
Flow matchingRegress the constant velocity along a straight noise-to-action path
Function evaluationsHow many times the expert is called per chunk; a straight path needs few
Action chunkA horizon of actions emitted at once and executed open-loop
Temporal ensemblingBlending overlapping predicted chunks to smooth the join between them
Latent world modelPredicts the next observation's representation so a policy can plan in imagination
Open-loop executionRunning a chunk without observing; long chunks react slowly, short ones cost more inference
7

Check your understanding

0/4 answered