Vision-language-action
A vision-language model reads a picture and ends its turn with a sentence. A robot has to end its turn with a movement instead, and a movement is not a word: it is a vector of joint targets or end-effector poses that has to be continuous, well-timed and physically possible. Joint targets are the angles each motor should hold, and an end-effector pose is where the gripper at the tip of the arm should be and how it should be turned, a position plus an orientation. Neither can be picked from a fixed list of words the way a caption can, which is the whole difficulty. A vision-language-action model — a VLA — closes that gap by treating an action as one more thing the trunk can emit, where the trunk is the shared sequence-processing network that already reads the image and the instruction. It emits that action sometimes as text tokens, and sometimes through a separate action expert trained with flow matching. An action here means one command to the robot, the action expert is a small extra network dedicated to producing it, and flow matching is the technique of carrying a sample along the shortest possible straight path from noise to a target. What makes the problem distinctive is that the model is graded on consequences rather than on the next token: a slightly wrong caption is a bad caption, and a slightly wrong action is a dropped cup. That single change of grading standard propagates through every design choice in this part, because where a captioning model may round off, a control model has to be precise. This part follows the action from a discretised token to a continuous flow, then out to the chunking and the world models that let a policy move before it has finished thinking. A policy is the component that maps what the model currently sees to an action; chunking means emitting several steps of action at once; and a world model predicts what will happen next, so the policy can plan against an imagined future rather than waiting for the real one.
From seeing to acting
A language model with a different output head
The shortest description of a vision-language-action model is this: it is the model from the earlier parts of this volume with the language head replaced, or augmented, by one that emits actions. The language head is the final layer that turns the trunk's internal representation into a distribution over the next word, and here it is either retrained to emit numbers or joined by a new head that does. The image tower patches the camera view, cutting it into tiles and turning each tile into a vector, the trunk reads those vectors together with the instruction, and instead of predicting the next word it predicts the next command. Everything the trunk learned about objects, spatial relations and instructions transfers, because recognising that the red block is left of the blue one, and that "pick up" names an object, is exactly what a manipulation instruction needs. That is why starting from a pretrained checkpoint matters so much: teaching a vision-language trunk the world from scratch is expensive, while fine-tuning an existing one on action data costs a fraction of it and keeps the knowledge. The new part is only the mapping from the trunk's representation to a control signal, meaning the numbers a controller turns into motor movement.
The output representation is where the design splits. One family keeps the language interface exactly as it is and discretises the action, meaning each continuous number is rounded onto a fixed grid of bins and represented by the bin's index, so an action is a short sequence of tokens the trunk can emit with the same cross-entropy loss it uses for words. Cross-entropy is the standard training objective that rewards putting probability on the correct next symbol, and it applies only when the target is one of a finite set, which the bins provide. The other family leaves the action continuous and adds a small expert that predicts it with a regression objective, meaning the expert outputs real numbers and is trained by shrinking the distance between its output and the target, usually a flow-matching velocity. The two routes are not merely stylistic: they make different assumptions about whether an action is naturally a category that can be named or a quantity that has to be measured. The diagram shows the shared pipeline; the two following steps take the two output routes in turn.
The shared pipeline: camera and instruction into the vision-language trunk, then an action output that is either discretised into tokens or produced by a continuous action expert.
The trunk is unchanged from the rest of the volume. A VLA is a multimodal model whose output head is a controller, which is why it inherits both the strengths and the failure modes of the underlying VLM.
Actions as tokens
Discretise a control signal and speak it
The first family borrows a trick from the tokenizer discussion: clip each dimension of a continuous action onto a fixed set of bins and emit the bin index as a token. Degrees of freedom are the independent axes a robot can move, so a seven-degree-of-freedom arm has seven of them, each driven by one number, and one action has seven dimensions. One action step is therefore seven tokens, one per axis, and a policy that wants to move for half a second emits a few dozen of them in a row, because commands are issued many times a second. The appeal is enormous — no new head, no new loss, no architectural change at all. The same trunk that can write "pick up the red block" can, after fine-tuning on action data, write the numbers that make it happen, where fine-tuning means continuing training on a smaller specialised set of robot demonstrations. It can also interleave reasoning text with the action tokens in one stream, so it can narrate what it is about to do and then do it.
The cost is quantisation, the rounding that replaces each exact value with the nearest bin's index and discards the difference. A bin represents a range of joint angles, so every emitted token carries an error of up to half a bin, and a controller that receives a coarse target has to smooth it or the motion stutters. The error is half a bin rather than a whole one because the token stands for the centre of its bin, so the worst a value can be off is half the bin's width. More bins reduce the error but lengthen the sequence, because every extra bin is another symbol the sampler can choose, and give the sampler more ways to emit an implausible combination; fewer bins are easy to predict but too coarse to be safe. Suppose a joint's range is divided into only sixteen bins: a command can then be off by a thirty-second of the joint's travel on either side, which at the end of a long reach becomes a visible wobble. The demo discretises a seven-dimensional action over an adjustable number of bins so the tension is visible lane by lane.
Seven action dimensions, each discretised into the number of bins on the slider. The highlighted cell per lane is the token the policy would emit; the trajectory step slider moves along a demonstration.
The flow-matching expert
A continuous head that predicts velocity, not a category
The second family refuses to discretise at all. It keeps the action continuous and attaches a small action expert to the pretrained trunk: a compact network that reads the trunk's representation of the scene and instruction and produces a velocity field over action space. A velocity field is a rule that gives a direction and a speed at any point, action space is the space of possible actions with one coordinate per joint, and the expert's job is to say which way and how fast to move through it. To feel the idea before its name, picture sliding a cart along straight rails at constant speed. Suppose the target action sets a gripper width of 0.4, and a noise sample starts it at 0.9. The simplest possible route from 0.9 to 0.4 is a straight line travelled at a constant speed, so a single number — the velocity — describes the entire journey, and if the expert predicts it correctly the sample lands exactly on the target. Training uses flow matching — a noise sample and a target action are connected by a straight line, and the expert is trained to predict the constant velocity that carries one to the other. Sampling starts at noise and integrates that velocity for a handful of steps, which is why these policies can run at control rates that a hundreds-step diffusion sampler cannot.
The reason the straight path matters is the number of function evaluations, meaning how many times the expert network has to be called to produce one action. A curved path needs many small steps to follow accurately, because each step approximates the curve only over a short distance; a path the model has been trained to make straight can be traversed in a few Euler steps, where an Euler step is the simplest way to advance along a direction. Because the path is straight and the velocity is constant along it, the expert only has to be right about one vector rather than about a schedule of changing directions, and there is no compounding of per-step error the way a curve invites. A few evaluations per action chunk is the difference between a policy that fits a control loop and one that does not, where the control loop is the observe-decide-act cycle the controller repeats many times a second. The plot below draws the oracle straight paths from a cloud of noise samples to one target action, and the slider walks along the path so you can read the velocity, which is constant by construction.
Straight flow paths from Diffusion.flowX carrying noise samples to a target action, with the constant velocity of Diffusion.flowVelocity. The slider walks the flow time from noise to action.
A straight path is a property the training objective encourages, not an assumption: regressing x1 − x0 along a line is what lets a few integration steps reach the action.
Chunking and world models
Move before the next thought finishes
A policy that emits one action per inference pays the full latency of a forward pass between every command. Inference is one run of the network, and latency is the wall-clock time that run takes. If the trunk takes fifty milliseconds and the controller runs at a hundred hertz, meaning a hundred commands per second, then the arm needs a fresh target every ten milliseconds while the model can supply one only every fifty, so the arm is idle most of the time. Action chunking fixes that by asking the model for a whole horizon of actions at once — the next sixteen or fifty steps — and executing them open-loop while the next inference runs in parallel. The horizon is the number of future steps requested, and open-loop means the arm keeps running the commands it was given without checking what actually happened. Chunking also smooths the motion, because a single inference fixes a consistent run of commands instead of re-deciding at every tick. The cost is reactivity: while the chunk is executed the policy cannot see a new observation, so a chunk that is too long responds slowly to a moved object and a chunk that is too short pays the inference tax on every step. Temporal ensembling, which blends overlapping chunks by averaging the predictions two successive inferences make for the same moment, is the usual compromise.
A latent world model addresses the other half of the loop. A world model is a network that predicts what happens next, and latent means it makes that prediction in the trunk's internal vector form rather than as a picture, which keeps the prediction cheap. If the trunk can predict the representation of the next observation as well as the action, then a policy can plan a short rollout in latent space, score candidate chunks against the predicted future, and act on the best one without touching the real environment. A rollout is that imagined sequence of steps, and scoring means comparing candidate actions against the future the model just predicted. The same prediction, trained on the same interleaved stream, tells the model what its action is likely to cause, which is what separates "emit a plausible movement" from "know what the movement will do". Planning against that predicted future is what lets a policy improve a chunk before it commits to it, which matters because a real arm cannot undo a movement once it is underway. The timeline below shows chunks being inferred, executed and re-planned as the horizon changes.
A control timeline: solid blocks are actions executed open-loop, the dashed extension is the part of the chunk a latent world model rolls forward, and the ticks are inference calls. The slider sets the chunk horizon in steps.
Longer chunks mean fewer inference calls and smoother motion, but slower reaction. The world model is what lets the policy spend the idle time imagining the future instead of waiting.
Where this shows up
From a sentence to a movement
Actions as language tokens
RT-2 discretises each action dimension and appends the tokens to the vocabulary, the model's fixed list of output symbols, so a web-scale vision-language model can be fine-tuned into a policy that still answers questions. A web-scale model is one pretrained on a very large and general mixture of text and images. The inheritance is the point: semantic understanding transfers with the tokens, so the robot carries a vocabulary of objects and relations it absorbed from ordinary text and images rather than from robot data alone.
A flow-matching action expert
The pi-zero family keeps a pretrained vision-language backbone and adds a flow-matching action expert, trained to produce smooth, high-frequency action chunks — commands issued many times a second — that the backbone alone could not emit, because a trunk trained to pick a word does not readily output a continuous number. It is the continuous route of step three, paired with the chunking of step four.
The full loop is where the rest of the guide meets control: the same latent world models that generate video (the sibling volume) are the ones a policy rolls forward, and the same token-budget arithmetic decides whether the camera view and the instruction fit beside the action horizon in one context. Token budget is the number of tokens a fixed context window can hold, and the camera view competes with the instruction for that room. A VLA sits on top of a control stack — the layered software between a command and a motor.
Further reading
These papers are the two output routes — discretised action tokens and a continuous action expert — together with the world-model idea the loop borrows from. Read together they trace one action from a word-like token, through a continuous velocity, to a chunk planned against an imagined future.
- Brianna Zitkovich, Tianhe Yu, Sichun Xu and coauthors, "RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control", 2023 — actions discretised into tokens in a vision-language vocabulary.
- Moo Jin Kim, Karl Pertsch, Siddharth Karamcheti and coauthors, "OpenVLA: An Open-Source Vision-Language-Action Model", 2024 — an open recipe for fine-tuning a VLM into a policy with discretised actions.
- Kevin Black, Noah Brown, Danny Driess and coauthors, "π0: A Vision-Language-Action Flow Model for General Robot Control", 2024 — a pretrained backbone with a flow-matching action expert producing action chunks.
- Danijar Hafner, Jurgis Pasukonis, Jimmy Ba and Timothy Lillicrap, "Mastering Diverse Domains through World Models", 2023 — DreamerV3, the latent world model that plans in a learned representation rather than in pixels.
- Anthony Brohan, Noah Brown, Justice Carbajal and coauthors, "RT-1: Robotics Transformer for Real-World Control at Scale", 2022 — the earlier token-based policy that made chunking and large-scale robot data concrete.
Cheat sheet
| Term | Meaning here |
|---|---|
| VLA | A multimodal trunk whose output head emits actions instead of, or beside, text |
| Discretised action | Each action dimension clipped onto bins, emitted as an ordinary token |
| Action token | One bin index; a seven-axis arm emits seven per action step |
| Bin count | Sets the worst-case command error and the length of the action sequence |
| Action expert | A small continuous head on the trunk, trained with regression rather than cross-entropy |
| Flow matching | Regress the constant velocity along a straight noise-to-action path |
| Function evaluations | How many times the expert is called per chunk; a straight path needs few |
| Action chunk | A horizon of actions emitted at once and executed open-loop |
| Temporal ensembling | Blending overlapping predicted chunks to smooth the join between them |
| Latent world model | Predicts the next observation's representation so a policy can plan in imagination |
| Open-loop execution | Running a chunk without observing; long chunks react slowly, short ones cost more inference |