Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

A computation graph

Every value is a node; every operation an edge

A derivative is a local statement: how does one number move when another moves? For a composition the data comes in pieces, and the cleanest way to hold them is a computation graph. Our running example is the smallest network that still has every ingredient. An input $x$ feeds an affine map, then a ReLU, then a second affine map, and finally a squared error against a target $t$:

$$ z_1 = w_1 x + b_1,\qquad a_1 = \mathrm{relu}(z_1),\qquad z_2 = w_2 a_1 + b_2,\qquad L = (z_2 - t)^2. $$

Each box below is one of those values; each arrow carries a number in the forward direction and, later, a sensitivity in the backward direction. Slide the five parameters and watch the forward values move along the graph. The whole representation is $2\times2$ weights plus biases, but nothing about the argument depends on the size.

The network as a graph. Node boxes hold the current forward values; edge labels hold the weights and the fixed operations.

2

The forward pass

Evaluate the composition one edge at a time

Running the network forward is just substituting left to right. The input and the biases are leaves: they have no incoming arrows and their values are given. A node with incoming arrows holds the value its parents produced. Step through the evaluation and the active set grows until the loss is reached — no calculus yet, only arithmetic, but every intermediate is being kept because the backward pass will need it.

Note the ReLU: $a_1 = \max(0, z_1)$. It is the one non-smooth piece, and everything about it is decided by the sign of $z_1$. When $z_1 \le 0$ the node is switched off, the forward value is zero, and — as the next step shows — so is its local derivative.

Visited nodes light up in evaluation order. Step advances one node; Play runs it on a timer.

3

Adjoints and the backward recursion

What must the output move for each node to move?

Fix a scalar output $L$ — here the loss — and ask for its sensitivity to every node. Define the adjoint of a node $v$ to be that sensitivity,

$$ \bar v \;=\; \frac{\partial L}{\partial v}. $$

For the output itself the adjoint is $1$. For any other node, $L$ reaches it only through the nodes that consume it, so the sensitivities add along every outgoing route. That is the chain rule written once per node:

$$ \bar v \;=\; \sum_{u \,\in\, \mathrm{children}(v)} \bar u \, \frac{\partial u}{\partial v}. $$

The factor $\partial u/\partial v$ is local: it depends only on the edge from $v$ to $u$ and on the forward values already stored. Multiplying it by the incoming adjoint $\bar u$ hands the sensitivity to $v$, and the recursion proceeds from the output back to the leaves. Step through it: each node shows its local derivative, the incoming adjoint, and their product. The magenta path traces one complete chain, say $\partial L/\partial x$.

Reverse sweep. The highlighted path is the chain rule for one derivative; each node reports local × incoming = product.

4

Checking one derivative against a finite difference

Agreement is the only honest test

Backprop returns a product of local factors, and a sign error in any one of them is invisible until training misbehaves. The independent check is the definition from The derivative: nudge the parameter and measure. For a scalar parameter $p$,

$$ \frac{\partial L}{\partial p} \;\approx\; \frac{L(p+h) - L(p-h)}{2h}, $$

which is exactly CalcViz.numericDeriv. Pick a parameter with the buttons and change the step $h$; the plot shows the loss curve, the backprop slope as a tangent, and the finite-difference secant. The two slopes are drawn from completely different information — one from the graph, one from re-running the forward pass twice — and they should agree to several digits. That agreement is what authorises the backward sweep.

Loss against the chosen parameter. Solid tangent slope = backprop; dashed secant slope = central finite difference over ±h.

5

One reverse sweep for every parameter

Why reverse mode is the right direction for a scalar loss

The check above costs two forward evaluations to differentiate one parameter. Training needs the gradient with respect to all of them, and the count explodes: central differences over $N$ parameters need $2N$ forward passes, so a million-parameter model would need two million passes for a single step. Backpropagation is independent of $N$ — one forward pass stores the values, one backward pass computes every adjoint — because each node's adjoint is reused by all of its parents.

$$ \text{forward-mode cost} \;\sim\; 2N \text{ passes},\qquad \text{reverse-mode cost} \;\sim\; 2 \text{ passes}. $$

This is the asymmetry that makes deep learning possible. Reverse mode is cheap exactly when the output is a scalar and the inputs are many, which is the shape of every training problem; forward mode wins in the mirror case of few inputs and many outputs. Slide the parameter count and watch the two lines separate on the log scale.

Forward passes needed for a full gradient against parameter count, on log–log axes. Finite differences (magenta) grow with N; one reverse sweep (blue) stays flat.

6

Where this shows up

The algorithm behind every training loop

Every weight update in a modern model is one term of this sweep. LLM Training, whose fifth part follows a loss backward through attention and the feed-forward blocks, is precisely this recursion applied to a larger graph. The single-variable version is The chain rule; the bookkeeping for whole matrices of parameters rather than single numbers is Matrix calculus, where the local factors become Jacobians and the products become matrix multiplies. Read together, those three pages are one idea at three scales.

7

Notation to carry forward

ObjectReads asWhere it is used
v, z, a, LNode values in the graph: pre-activation, activation, scalar lossStep 1, Step 2
∂u/∂vLocal derivative on the edge v → u; fixed by the operationStep 3, Step 4
v̄ = ∂L/∂vAdjoint: sensitivity of the scalar output to a nodeStep 3
v̄ = Σᵤ ū · ∂u/∂vThe backward recursion — the chain rule summed over consumersStep 3
2N vs 2Forward passes for a full finite-difference gradient versus one reverse sweepStep 5
8

Further reading

9

Check your understanding

0/3 answered