The chain rule
A composition is two machines bolted together: the first takes the input and produces an intermediate shaft angle, the second turns that angle into the output. Each machine has its own rate gauge, and the chain rule says the overall rate is simply the two gauges multiplied. Drag the input and watch the product track the real rate.
A machine made of machines
Composition, and the difference quotient that multiplies
Our robot on a line now moves through two stages. The motor's clock $t$ drives an intermediate shaft whose angle is $u = g(t)$; a nonlinear linkage turns that angle into the robot's position $y = f(u)$. The position is the composition of the two machines:
Let $t$ move by a small amount $\Delta t$. That moves $u$ by $\Delta u$, which moves $y$ by $\Delta y$. The three changes are tied together by an exact identity for any nonzero steps — the $\Delta u$ cancels like an ordinary fraction:
Now let $\Delta t \to 0$. A differentiable $g$ is continuous, so $\Delta u \to 0$ too, and both factors slide onto derivatives. That is the chain rule:
Read it as an instruction: differentiate the outer function at the inner value, then multiply by the inner function's derivative. The gear train below is the picture. The inner shaft rate is $\tfrac{du}{dt} = kt$ and the outer linkage rate is $\tfrac{dy}{du} = \cos u$; their product is the output rate. Slide $t$ and the inner gain $k$ and watch the two gauges and their product move together.
Two meshing gears: the inner gear turns at $\tfrac{du}{dt}$, the outer at $\tfrac{dy}{du}$. The bars below are rate gauges; the output rate is their product.
Compose two functions
The rule at a point, checked numerically
The identity is general. Pick an inner function $g$ and an outer function $f$ and the graph below draws $y = f(g(x))$. Drag the probe, or use the slider, and the readout compares two numbers: the derivative computed directly from the composed curve, and the product $f'(g(x))\,g'(x)$ assembled from the two pieces. They agree because the cancellation above is exact before the limit.
The composed curve $f(g(x))$ with its tangent at the probe. Drag the probe point along the curve.
Rates you can multiply
Why the units work out
Chain-rule factors are rates with units, and the units cancel the same way the $\Delta u$ did. An engine shaft turns at $\omega_e$ revolutions per second. A gearbox with ratio $R$ slows the wheel to $\omega_w = \omega_e / R$, so $\tfrac{d\omega_w}{d\omega_e} = \tfrac{1}{R}$ — revolutions of wheel per revolution of engine, dimensionless. The wheel of radius $r$ then rolls out $\tfrac{dv}{d\omega_w} = 2\pi r$ metres of ground per revolution. Multiply the factors and the output rate is
Change the gear ratio and the output rate scales by exactly the same factor. Slide any of the three inputs; the speedometer and the readout recompute the chain.
Engine gear meshing with the wheel, plus a speedometer for the rolling rate $v = \omega_w \cdot 2\pi r$.
Running the chain outward-in
A preview of backpropagation
When a composition has more than two layers, the chain rule just keeps multiplying. Here three stages turn $t$ into $u_1$, $u_1$ into $u_2$, and $u_2$ into the output $y$:
To differentiate, you sweep from the outside in: start at the output, read off the local derivative on each edge, and accumulate the product as you travel back toward the input. Press play to run the sweep, and slide $t$ to see each local factor change. This is the whole idea backpropagation needs — the graph is just wider, and the multiplication runs in reverse.
A three-layer chain. The highlighted edges are the factors already picked up by the outward-in sweep; their product is the running gradient.
| Layer | Local derivative | Value | Running product |
|---|
Where this shows up
The rule the whole site runs on
Backpropagation is this rule evaluated in reverse on a computation graph. Each edge carries a local derivative; the gradient of a loss with respect to any weight is the product of the factors along the path from that weight out to the loss. That is exactly Volume II, Part 9, where the chain rule is applied layer by layer through a small network. The same multiplication is why training a language model is tractable at all: every gradient in LLM training is a long chain of local rates, swept outward-in. And running the rule the other direction — integrating the factors — is substitution, in Part 10.
The chain rule on one card
| Form | Reads as | Where it comes up |
|---|---|---|
(f ∘ g)′(x) = f′(g(x)) g′(x) | Outer derivative at the inner value, times inner derivative | Every composition, this part |
dy/dx = dy/du · du/dx | The rate is the product of the two stage rates | Related rates (Part 6), units |
Δy/Δx = (Δy/Δu)(Δu/Δx) | The exact finite-step identity the limit is taken from | The derivation, Step 1 |
dy/dx = dy/du₁ · du₁/du₂ · … · duₙ/dx | Keep multiplying for a chain of any depth | Backpropagation (Vol II Part 9) |
d/dx f(g(x)) | Differentiate outside-in; never forget the inner factor | The most common mistake in the subject |
Further reading
- 3Blue1Brown, Visualizing the chain rule and product rule — the gears-and-rates picture, animated.
- MIT 18.01, Single Variable Calculus — the rigorous derivation and plenty of composition drills.
- Khan Academy, Chain rule introduction — worked examples if the mechanics still feel slippery.