Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

Gradients of a scalar

One partial derivative per entry

Everything on this page is built from one object: a single scalar output, $f$, that depends on many inputs. If those inputs are the entries of a vector $x \in \mathbb{R}^n$, the gradient is the vector of partial derivatives, with the same shape as $x$:

$$ \nabla_x f \;=\; \begin{bmatrix} \partial f / \partial x_1 \\ \vdots \\ \partial f / \partial x_n \end{bmatrix} \in \mathbb{R}^{n \times 1}, \qquad (\nabla_x f)_i \;=\; \frac{\partial f}{\partial x_i}. $$

If the inputs are the entries of a matrix $A \in \mathbb{R}^{m \times n}$, the gradient is a matrix of the same shape:

$$ (\nabla_A f)_{ij} \;=\; \frac{\partial f}{\partial A_{ij}}. $$

That is the whole definition. Differentiating with respect to a matrix is not a new kind of limit — it is $m n$ ordinary partial derivatives, stacked back into the shape you started with. Our running example stays the two-link arm reaching for a target: its Jacobian (Part 5) already measured how the fingertip moves when the joint angles change, and its Hessian will stack second derivatives the same way. A small two-layer net joins the picture in Step 5.

The only genuine subtlety is arrangement — which dimension goes first when the derivative is not square. That is the layout question, and it is worth getting straight before any identity.

⚠️ Two careless assumptions cause most matrix-calculus bugs: that a gradient has the same shape as whatever you "differentiated by", and that transposes can be moved around freely. Both are convention-dependent, and the convention that makes the first true is not the one that makes the second obvious.
2

Shape checker

Count dimensions before computing anything

Pick an expression and set the sizes $m \times n$. The canvas shows what the forward pass produced and what shape the derivative has under the numerator layout — output dimensions first, input dimensions second. When you can predict these two boxes, you have avoided the most common error before it happens.

Forward shape (middle) and numerator-layout derivative shape (right), for the expression and sizes you choose.

Square-only expressions ($x^{\top} A x$, $\log\det A$) use the smaller side for the square matrix, so the picture never asks for a shape that does not exist.
3

Two layouts, one transpose apart

Numerator first, or denominator first

For a vector-to-vector map $y = f(x)$ with $y \in \mathbb{R}^m$ and $x \in \mathbb{R}^n$, the Jacobian is an $m \times n$ table of partials. Nothing forces you to store it either way round, so two conventions grew up:

$$ \underbrace{\frac{\partial y}{\partial x}}_{\text{numerator layout}} \in \mathbb{R}^{m \times n}, \qquad \underbrace{\frac{\partial y}{\partial x}}_{\text{denominator layout}} \in \mathbb{R}^{n \times m}. $$

Numerator layout keeps the output first; denominator layout keeps the input first. They hold exactly the same numbers — one is the transpose of the other. Toggle the button and watch the grid flip. The slider changes one entry of $A$ so you can see the transpose relation is entrywise: $(J^{\top})_{ji} = J_{ij}$.

The Jacobian of $y = A x$ in both layouts. The active convention is drawn opaque; the other is dimmed.

💡 Why ML code uses denominator layout. A learning rule wants to subtract the gradient from the parameter: $W \leftarrow W - \eta\,\partial L/\partial W$. That line only type-checks if $\partial L/\partial W$ has the same shape as $W$ — denominator layout. PyTorch's .backward() therefore hands you a gradient shaped like the tensor you differentiated against; the transpose is hidden in the storage order, not in your code.
4

The core identities, by indices

Derive, then verify against finite differences

Each identity below comes from writing the scalar out as a sum and differentiating one entry at a time. For the quadratic form, $f(x) = x^{\top} A x = \sum_{i,j} x_i A_{ij} x_j$, so

$$ \frac{\partial f}{\partial x_k} \;=\; \sum_j A_{kj} x_j + \sum_i x_i A_{ik} \;=\; (Ax)_k + (A^{\top}x)_k \quad\Longrightarrow\quad \nabla_x\, x^{\top} A x \;=\; (A + A^{\top}) x. $$

For the trace, $\operatorname{tr}(AB) = \sum_{i,j} A_{ij} B_{ji}$, so $\partial/\partial A_{kl}$ picks out $B_{lk}$ and the gradient is $B^{\top}$. For the log-determinant, Jacobi's formula gives $\partial \log\det A / \partial A_{ij} = (A^{-1})_{ji}$, i.e. $\nabla_A \log\det A = A^{-\top}$. The button verifies one identity at a time: the analytic gradient is compared entry-by-entry with a central finite difference, and the readout reports the largest absolute mismatch.

Left: the analytic gradient. Right: the same gradient from finite differences. The worst-mismatch cell is outlined.

Slide $h$ smaller and the error usually falls, then rises again as floating-point cancellation takes over — the same trade-off you met in Step 2 of the derivative page.
5

The chain rule for a scalar loss

Backprop in one layer, previewed

The reason these identities matter is that a training loss composes them. Take a tiny net with affine map $a = W x + b$, a ReLU $h = \operatorname{relu}(a)$, and a linear read-out $y = w^{\top} h + c$. With a squared loss $L = \tfrac{1}{2}(y - t)^2$, the chain rule runs backwards through the shapes:

$$ \frac{\partial L}{\partial y} = (y - t), \qquad \frac{\partial L}{\partial h} = \frac{\partial L}{\partial y}\, w, \qquad \frac{\partial L}{\partial a_k} = \frac{\partial L}{\partial h_k}\,\mathbb{1}[a_k > 0], \qquad \frac{\partial L}{\partial W} = \frac{\partial L}{\partial a}\, x^{\top}. $$

That last step is an outer product: a column of length (hidden) times a row of length (input) gives a matrix shaped like $W$ — denominator layout again. Drag the entry $W_{00}$ and watch the loss, the analytic gradient, and the finite-difference gradient move together.

Every entry of $\partial L/\partial W$: bars are matrix calculus, dots are finite differences. The slider entry is outlined.

6

Reference card

A lookup table you can keep open

Pick an identity to see its formula, its shapes, and a diagram of the objects. This is the whole page compressed into one widget; the cheat table further down adds the notation.

Input shape, forward shape, and gradient shape for the selected identity.

$$ \frac{d}{dx}\,(Ax) = A^{\top} $$
7

Where this shows up

The identities are the load-bearing walls

Training a language model is this page applied at scale: every weight update is $\partial L/\partial W$ in denominator layout, and the chain rule of Step 5 is exactly what LLM Training automates over billions of parameters. The next part, Backpropagation is the chain rule, turns the sequence of identities into an algorithm on a computation graph. And the Gauss–Newton step in Nonlinear Optimization, $H = J^{\top} J$, is matrix calculus twice: the Jacobian $J = \partial r/\partial \theta$ from Part 5, and this part's transpose and product rules composed into a normal matrix.

8

Notation to carry forward

ExpressionMeaningShape (denominator layout)
∂f/∂xGradient of a scalar with respect to a vectorsame as x
∂f/∂XGradient of a scalar with respect to a matrixsame as X
∂y/∂xJacobian of a vector map; rows follow the input in denominator layout(dim x) × (dim y)
∂(Ax)/∂x = AᵀLinear map derivativesame as the gradient of x
∂(xᵀAx)/∂x = (A+Aᵀ)xQuadratic form gradient (symmetric if A is)as x
∂tr(AB)/∂A = BᵀTrace under a matrix perturbationas A
∂log det A/∂A = A⁻ᵀLog-determinant (Jacobi's formula); needs A invertibleas A
9

Further reading

10

Check your understanding

0/3 answered