Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

The two path families

A path from noise to data, and how straight it is

Every generative model that walks from noise to data needs a bridge: a curve that starts at a noise sample $x_0$ and ends at a data sample $x_1$. The simplest one is a straight line. With a path parameter $t$ running from $0$ to $1$, the point on the line is $x_t = (1-t)\,x_0 + t\,x_1$, which is exactly what Diffusion.flowX computes. The curve on the top canvas is a different bridge with the same two endpoints, $x_t = (1-t)x_0 + t x_1 + b\sin(\pi t)$, bowed upward by the constant $b$ — the kind of path a variance-preserving diffusion model actually induces.

The dashed polylines are what a sampler can build with a finite number of Euler steps. The straight bridge is reproduced exactly by a single step, because the velocity never changes direction along it. The curved bridge is not: each Euler step follows the tangent at the current point and drifts off the curve, and with few steps the surviving error is visible. Drag the slider and watch the two error curves below.

Top: the two bridges between the same endpoints, with tangent arrows for each velocity field and dashed Euler polylines at the current step count. Bottom: endpoint error against the number of Euler steps, on a log scale.

The straight path is exact at $N=1$; the curved one keeps improving as $N$ grows. Substituting a straight target is what buys low-step sampling.

💡 By the end of this part you'll be able to read a flow-matching paper as a statement about a velocity field on a path, explain why the training target is just $x_1 - x_0$, and say why a straight interpolation needs far fewer sampling steps than a curved one.
2

The velocity field and the ODE

The network is a field, the sampler is an integrator

The flow-matching view rewrites generation as a deterministic ordinary differential equation. There is a velocity field $v_\theta(x, t)$ defined at every point and time, and a sample follows it by integrating

$$ \frac{dx}{dt} = v_\theta(x, t), \qquad x(0) = x_0 \sim \mathcal{N}(0, I). $$

Nothing here is stochastic once the starting noise is fixed. The randomness that a diffusion sampler injects at every ancestral step is gone; what remains is a single trajectory through a vector field, determined by where it started. That is a real simplification, because ordinary differential equations come with a mature numerical toolkit. Euler's method is the crudest: take a step in the direction the field points right now, $x \leftarrow x + v_\theta(x, t)\,\Delta t$, and repeat. Higher-order integrators use a few field evaluations per step to follow curvature more accurately.

The arrows on the top canvas are the field sampled along each path. Along the straight bridge the arrows all point in the same direction and have the same length — the field is constant, which is precisely why one Euler step finishes the job. Along the curved bridge the arrows rotate, and no single step can capture the rotation. The forward direction is easy to state and the reverse is the hard part, exactly as in diffusion: $v_\theta$ is what has to be learned.

⚠ An ODE is not automatically a good sampler. Fewer stochastic terms makes the map deterministic and reproducible, but accuracy still depends on the field being easy to integrate. A badly curved field with a coarse step size gives a trajectory that misses the data manifold, and no amount of seed-fixing repairs it.
3

The training target

Regress $x_1 - x_0$

To train $v_\theta$ you need a target velocity at each point of the path, and for the straight bridge the answer is immediate. Differentiate $x_t = (1-t)x_0 + t x_1$ with respect to $t$ and the dependence on $t$ disappears:

$$ v = \frac{dx_t}{dt} = x_1 - x_0. $$

The target is the difference between the two endpoints, constant along the path. This is conditional flow matching: condition on a particular pair $(x_0, x_1)$, regress the network's velocity at $x_t$ toward that constant, and average the squared error over pairs. In code it is the one-liner Diffusion.flowLoss(vPred, x0, x1), which returns $(v_\theta(x_t,t) - (x_1 - x_0))^2$. Sampling a training point is cheap: draw $x_1$ from the data, draw $x_0$ from the noise, draw $t$ uniformly, form $x_t$ on the line, and regress. There is no chain to unroll and no schedule to tune.

The conditional objectives average out to the right thing. If the network minimises the error against $x_1 - x_0$ at every point, its prediction converges to the conditional expectation of that difference given the current position, which is the marginal velocity field that transports the noise distribution onto the data distribution. The curved family needs the tangential derivative instead, Diffusion.curvedVelocity, which changes with $t$ and forces the network to represent rotation as well as translation.

💡 The target $x_1 - x_0$ is why flow matching is so easy to implement: no signal-to-noise bookkeeping, no reparameterised $\epsilon$ or $v$, just a difference of two samples. The schedule is replaced by the geometry of the path.
4

Why straighter is cheaper

One step, four steps, fifty

Look again at the bottom canvas. The straight-path error sits on the floor at every $N$, because a constant field is integrated exactly by Euler at any step size. The curved-path error falls as the steps multiply, and it never reaches the floor until the steps are dense enough that the polyline is indistinguishable from the curve. A sampler that must resolve curvature pays for it in network evaluations, and each evaluation is a forward pass through a large model.

Rectified flow turns that observation into an algorithm. Train the field on a set of pairs, then use the trained field to trace new pairs through the model — the noise that maps to a given data point under the current field — and retrain on those straighter pairings. Iterating this reflow step straightens the probability-flow trajectories, and as they straighten, the number of Euler steps needed to follow them collapses. Enough reflow passes and one or two steps suffice, which is what makes single-step and few-step image generators possible. The same argument motivates consistency and distillation methods, which also trade training compute for sampling compute.

⚠ Fewer steps is not a free lunch. Straightening is bought with extra training rounds, and a field that is easy to integrate is not automatically the field that reproduces the data distribution most faithfully. Very low-step samplers trade sample diversity for speed, the same tension that guidance introduces in the next part.
5

Where this shows up

The frontier trainers, rewritten

Image

Stable Diffusion 3 and Flux

Both replace the DDPM schedule with a rectified-flow objective over a straight interpolation between noise and latent. The latent diffusion machinery is unchanged; only the target the network regresses changes. SD3 reports that the straight-path objective lets it sample well at far fewer steps than the equivalent diffusion model.

Control & robotics

π₀ and action chunking

Robot policies run the same trick on a different modality: a flow-matching head predicts a velocity field over actions rather than pixels, and the policy integrates it in a handful of steps at control rate. The straight-path target is what makes a real-time budget feasible, and the sampling chapter is where the step-count trade is made explicit.

The pattern generalises to video, audio and any continuous signal: replace the noise schedule with a path, regress the velocity, and integrate. The ODE guide develops the numerical side of this picture — Euler, Runge–Kutta and their error orders — which is exactly the toolkit a flow sampler uses.

Further reading

Flow matching grew out of an older idea — continuous normalising flows — that was rescued by a change of objective. These papers are the ones to read, in roughly this order, and the schedule-to-path rewrite in the third is the clearest statement of why the field changed.

Lipman and coauthors introduce the conditional-flow-matching objective and prove the averaging argument; Liu and coauthors give the rectified-flow algorithm and the reflow straightening step; Albergo and Vanden-Eijnden place both inside the stochastic-interpolant framework; and Esser and coauthors show the objective at production scale in SD3.

Cheat sheet

TermMeaning here
$x_t = (1-t)x_0 + t x_1$The straight bridge from noise to data; Diffusion.flowX
$v = x_1 - x_0$The conditional target velocity; constant along the straight path
ODE view$\dfrac{dx}{dt} = v_\theta(x,t)$; sampling is numerical integration
Euler step$x \leftarrow x + v_\theta(x,t)\,\Delta t$; exact only for a constant field
Curved pathSame endpoints, rotating tangent; needs many steps to integrate accurately
ReflowRetrain on pairs traced through the current field, straightening the paths
Few-step sampling1–4 Euler steps once trajectories are close to straight
Adopted bySD3, Flux, π₀ and most current frontier trainers
7

Check your understanding

0/4 answered