Flow matching and rectified flow
The reverse process of a diffusion model is an ordinary differential equation, and the network that drives it is a velocity field. Flow matching trains that field directly: draw a noise sample, draw a data sample, connect them, and regress the velocity along the connecting path. The choice of path is not innocent — a straight path integrates in a handful of Euler steps while a curved one needs many — and that single observation is why the frontier trainers replaced the older noise schedules with flow matching and rectified flow.
The two path families
A path from noise to data, and how straight it is
Every generative model that walks from noise to data needs a bridge: a curve that starts at a noise sample $x_0$ and ends at a data sample $x_1$. The simplest one is a straight line. With a path parameter $t$ running from $0$ to $1$, the point on the line is $x_t = (1-t)\,x_0 + t\,x_1$, which is exactly what Diffusion.flowX computes. The curve on the top canvas is a different bridge with the same two endpoints, $x_t = (1-t)x_0 + t x_1 + b\sin(\pi t)$, bowed upward by the constant $b$ — the kind of path a variance-preserving diffusion model actually induces.
The dashed polylines are what a sampler can build with a finite number of Euler steps. The straight bridge is reproduced exactly by a single step, because the velocity never changes direction along it. The curved bridge is not: each Euler step follows the tangent at the current point and drifts off the curve, and with few steps the surviving error is visible. Drag the slider and watch the two error curves below.
Top: the two bridges between the same endpoints, with tangent arrows for each velocity field and dashed Euler polylines at the current step count. Bottom: endpoint error against the number of Euler steps, on a log scale.
The straight path is exact at $N=1$; the curved one keeps improving as $N$ grows. Substituting a straight target is what buys low-step sampling.
The velocity field and the ODE
The network is a field, the sampler is an integrator
The flow-matching view rewrites generation as a deterministic ordinary differential equation. There is a velocity field $v_\theta(x, t)$ defined at every point and time, and a sample follows it by integrating
$$ \frac{dx}{dt} = v_\theta(x, t), \qquad x(0) = x_0 \sim \mathcal{N}(0, I). $$
Nothing here is stochastic once the starting noise is fixed. The randomness that a diffusion sampler injects at every ancestral step is gone; what remains is a single trajectory through a vector field, determined by where it started. That is a real simplification, because ordinary differential equations come with a mature numerical toolkit. Euler's method is the crudest: take a step in the direction the field points right now, $x \leftarrow x + v_\theta(x, t)\,\Delta t$, and repeat. Higher-order integrators use a few field evaluations per step to follow curvature more accurately.
The arrows on the top canvas are the field sampled along each path. Along the straight bridge the arrows all point in the same direction and have the same length — the field is constant, which is precisely why one Euler step finishes the job. Along the curved bridge the arrows rotate, and no single step can capture the rotation. The forward direction is easy to state and the reverse is the hard part, exactly as in diffusion: $v_\theta$ is what has to be learned.
The training target
Regress $x_1 - x_0$
To train $v_\theta$ you need a target velocity at each point of the path, and for the straight bridge the answer is immediate. Differentiate $x_t = (1-t)x_0 + t x_1$ with respect to $t$ and the dependence on $t$ disappears:
$$ v = \frac{dx_t}{dt} = x_1 - x_0. $$
The target is the difference between the two endpoints, constant along the path. This is conditional flow matching: condition on a particular pair $(x_0, x_1)$, regress the network's velocity at $x_t$ toward that constant, and average the squared error over pairs. In code it is the one-liner Diffusion.flowLoss(vPred, x0, x1), which returns $(v_\theta(x_t,t) - (x_1 - x_0))^2$. Sampling a training point is cheap: draw $x_1$ from the data, draw $x_0$ from the noise, draw $t$ uniformly, form $x_t$ on the line, and regress. There is no chain to unroll and no schedule to tune.
The conditional objectives average out to the right thing. If the network minimises the error against $x_1 - x_0$ at every point, its prediction converges to the conditional expectation of that difference given the current position, which is the marginal velocity field that transports the noise distribution onto the data distribution. The curved family needs the tangential derivative instead, Diffusion.curvedVelocity, which changes with $t$ and forces the network to represent rotation as well as translation.
Why straighter is cheaper
One step, four steps, fifty
Look again at the bottom canvas. The straight-path error sits on the floor at every $N$, because a constant field is integrated exactly by Euler at any step size. The curved-path error falls as the steps multiply, and it never reaches the floor until the steps are dense enough that the polyline is indistinguishable from the curve. A sampler that must resolve curvature pays for it in network evaluations, and each evaluation is a forward pass through a large model.
Rectified flow turns that observation into an algorithm. Train the field on a set of pairs, then use the trained field to trace new pairs through the model — the noise that maps to a given data point under the current field — and retrain on those straighter pairings. Iterating this reflow step straightens the probability-flow trajectories, and as they straighten, the number of Euler steps needed to follow them collapses. Enough reflow passes and one or two steps suffice, which is what makes single-step and few-step image generators possible. The same argument motivates consistency and distillation methods, which also trade training compute for sampling compute.
Where this shows up
The frontier trainers, rewritten
Stable Diffusion 3 and Flux
Both replace the DDPM schedule with a rectified-flow objective over a straight interpolation between noise and latent. The latent diffusion machinery is unchanged; only the target the network regresses changes. SD3 reports that the straight-path objective lets it sample well at far fewer steps than the equivalent diffusion model.
π₀ and action chunking
Robot policies run the same trick on a different modality: a flow-matching head predicts a velocity field over actions rather than pixels, and the policy integrates it in a handful of steps at control rate. The straight-path target is what makes a real-time budget feasible, and the sampling chapter is where the step-count trade is made explicit.
The pattern generalises to video, audio and any continuous signal: replace the noise schedule with a path, regress the velocity, and integrate. The ODE guide develops the numerical side of this picture — Euler, Runge–Kutta and their error orders — which is exactly the toolkit a flow sampler uses.
Further reading
Flow matching grew out of an older idea — continuous normalising flows — that was rescued by a change of objective. These papers are the ones to read, in roughly this order, and the schedule-to-path rewrite in the third is the clearest statement of why the field changed.
Lipman and coauthors introduce the conditional-flow-matching objective and prove the averaging argument; Liu and coauthors give the rectified-flow algorithm and the reflow straightening step; Albergo and Vanden-Eijnden place both inside the stochastic-interpolant framework; and Esser and coauthors show the objective at production scale in SD3.
- Yaron Lipman, Ricky Chen, Heli Ben-Hamu, Maximilian Nickel and Matt Le, "Flow Matching for Generative Modeling", 2022 — the conditional objective and the velocity-field view.
- Xingchao Liu, Chengyue Gong and Qiang Liu, "Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow", 2022 — reflow and the straightening argument.
- Michael Albergo and Eric Vanden-Eijnden, "Building Normalizing Flows with Stochastic Interpolants", 2023 — paths, velocities and the general interpolation framework.
- Patrick Esser and coauthors, "Scaling Rectified Flow Transformers for High-Resolution Image Synthesis", 2024 — SD3, and the case for rectified flow over the older schedule.
Cheat sheet
| Term | Meaning here |
|---|---|
| $x_t = (1-t)x_0 + t x_1$ | The straight bridge from noise to data; Diffusion.flowX |
| $v = x_1 - x_0$ | The conditional target velocity; constant along the straight path |
| ODE view | $\dfrac{dx}{dt} = v_\theta(x,t)$; sampling is numerical integration |
| Euler step | $x \leftarrow x + v_\theta(x,t)\,\Delta t$; exact only for a constant field |
| Curved path | Same endpoints, rotating tangent; needs many steps to integrate accurately |
| Reflow | Retrain on pairs traced through the current field, straightening the paths |
| Few-step sampling | 1–4 Euler steps once trajectories are close to straight |
| Adopted by | SD3, Flux, π₀ and most current frontier trainers |