Matrix calculus
So far the derivative has been a number, then a slope function, then — for a map between spaces — a matrix. This part makes that matrix idea official: how to differentiate a scalar with respect to a vector or a matrix, the two conventions that disagree only by a transpose, and the half-dozen identities that do almost all the work when you train a network.
Gradients of a scalar
One partial derivative per entry
Everything on this page is built from one object: a single scalar output, $f$, that depends on many inputs. If those inputs are the entries of a vector $x \in \mathbb{R}^n$, the gradient is the vector of partial derivatives, with the same shape as $x$:
If the inputs are the entries of a matrix $A \in \mathbb{R}^{m \times n}$, the gradient is a matrix of the same shape:
That is the whole definition. Differentiating with respect to a matrix is not a new kind of limit — it is $m n$ ordinary partial derivatives, stacked back into the shape you started with. Our running example stays the two-link arm reaching for a target: its Jacobian (Part 5) already measured how the fingertip moves when the joint angles change, and its Hessian will stack second derivatives the same way. A small two-layer net joins the picture in Step 5.
The only genuine subtlety is arrangement — which dimension goes first when the derivative is not square. That is the layout question, and it is worth getting straight before any identity.
Shape checker
Count dimensions before computing anything
Pick an expression and set the sizes $m \times n$. The canvas shows what the forward pass produced and what shape the derivative has under the numerator layout — output dimensions first, input dimensions second. When you can predict these two boxes, you have avoided the most common error before it happens.
Forward shape (middle) and numerator-layout derivative shape (right), for the expression and sizes you choose.
Two layouts, one transpose apart
Numerator first, or denominator first
For a vector-to-vector map $y = f(x)$ with $y \in \mathbb{R}^m$ and $x \in \mathbb{R}^n$, the Jacobian is an $m \times n$ table of partials. Nothing forces you to store it either way round, so two conventions grew up:
Numerator layout keeps the output first; denominator layout keeps the input first. They hold exactly the same numbers — one is the transpose of the other. Toggle the button and watch the grid flip. The slider changes one entry of $A$ so you can see the transpose relation is entrywise: $(J^{\top})_{ji} = J_{ij}$.
The Jacobian of $y = A x$ in both layouts. The active convention is drawn opaque; the other is dimmed.
.backward() therefore hands you a gradient shaped like the tensor you differentiated against; the transpose is hidden in the storage order, not in your code.The core identities, by indices
Derive, then verify against finite differences
Each identity below comes from writing the scalar out as a sum and differentiating one entry at a time. For the quadratic form, $f(x) = x^{\top} A x = \sum_{i,j} x_i A_{ij} x_j$, so
For the trace, $\operatorname{tr}(AB) = \sum_{i,j} A_{ij} B_{ji}$, so $\partial/\partial A_{kl}$ picks out $B_{lk}$ and the gradient is $B^{\top}$. For the log-determinant, Jacobi's formula gives $\partial \log\det A / \partial A_{ij} = (A^{-1})_{ji}$, i.e. $\nabla_A \log\det A = A^{-\top}$. The button verifies one identity at a time: the analytic gradient is compared entry-by-entry with a central finite difference, and the readout reports the largest absolute mismatch.
Left: the analytic gradient. Right: the same gradient from finite differences. The worst-mismatch cell is outlined.
The chain rule for a scalar loss
Backprop in one layer, previewed
The reason these identities matter is that a training loss composes them. Take a tiny net with affine map $a = W x + b$, a ReLU $h = \operatorname{relu}(a)$, and a linear read-out $y = w^{\top} h + c$. With a squared loss $L = \tfrac{1}{2}(y - t)^2$, the chain rule runs backwards through the shapes:
That last step is an outer product: a column of length (hidden) times a row of length (input) gives a matrix shaped like $W$ — denominator layout again. Drag the entry $W_{00}$ and watch the loss, the analytic gradient, and the finite-difference gradient move together.
Every entry of $\partial L/\partial W$: bars are matrix calculus, dots are finite differences. The slider entry is outlined.
Reference card
A lookup table you can keep open
Pick an identity to see its formula, its shapes, and a diagram of the objects. This is the whole page compressed into one widget; the cheat table further down adds the notation.
Input shape, forward shape, and gradient shape for the selected identity.
Where this shows up
The identities are the load-bearing walls
Training a language model is this page applied at scale: every weight update is $\partial L/\partial W$ in denominator layout, and the chain rule of Step 5 is exactly what LLM Training automates over billions of parameters. The next part, Backpropagation is the chain rule, turns the sequence of identities into an algorithm on a computation graph. And the Gauss–Newton step in Nonlinear Optimization, $H = J^{\top} J$, is matrix calculus twice: the Jacobian $J = \partial r/\partial \theta$ from Part 5, and this part's transpose and product rules composed into a normal matrix.
Notation to carry forward
| Expression | Meaning | Shape (denominator layout) |
|---|---|---|
∂f/∂x | Gradient of a scalar with respect to a vector | same as x |
∂f/∂X | Gradient of a scalar with respect to a matrix | same as X |
∂y/∂x | Jacobian of a vector map; rows follow the input in denominator layout | (dim x) × (dim y) |
∂(Ax)/∂x = Aᵀ | Linear map derivative | same as the gradient of x |
∂(xᵀAx)/∂x = (A+Aᵀ)x | Quadratic form gradient (symmetric if A is) | as x |
∂tr(AB)/∂A = Bᵀ | Trace under a matrix perturbation | as A |
∂log det A/∂A = A⁻ᵀ | Log-determinant (Jacobi's formula); needs A invertible | as A |
Further reading
- The Matrix Cookbook — the exhaustive identity reference; the four identities above are its opening section.
- Wikipedia, Matrix calculus — a careful side-by-side of the numerator and denominator conventions.
- PyTorch autograd tutorial — the fan-in/fan-out view that turns these identities into backpropagation.