The derivative as a matrix
Last part, a derivative was a slope. Now the thing being differentiated is itself a vector: a two-link arm's tip is a function of two joint angles, and its derivative is a small matrix — the Jacobian — that turns joint speeds into tip speed. Differentiate that and you get the Hessian, a symmetric matrix whose eigenvalues decide whether a point is a bowl, a ridge or a saddle.
The Jacobian of the arm
A vector map differentiated is a matrix
Our running example is a two-link planar arm. Its end-effector position is a vector-valued function of the joint angles,
Each output component gets its own derivative with respect to each input. Stacking them gives the Jacobian, the best linear approximation of $p$ near the current configuration:
This is Part 12's gradient, done once per output. The columns are especially readable: they are the velocity of the tip when only one joint moves. Drag the angles and watch the two arrows at the end-effector — they are those columns. Then set joint rates $\theta'$ and read the tip velocity $p' = J\theta'$.
Arm geometry with the two Jacobian columns drawn from the tip (blue = $\partial p/\partial\theta_1$, pink = $\partial p/\partial\theta_2$) and the actual tip velocity $p'=J\theta'$ in dark teal.
What J does to a neighbourhood
A square becomes a parallelogram; area scales by det J
Because $J$ is the local linear model, a tiny square of joint perturbations maps to a parallelogram of tip displacements. Left panel: a square of half-width $h$ in the $(\theta_1,\theta_2)$ plane. Right panel: its image under $J$, a sheared parallelogram — the columns of $J$ are its two edge directions.
The determinant is the area-scaling factor. Differentiating the arm formula gives the clean identity
which vanishes at $\theta_2 = 0$ or $\theta_2 = \pi$: the fully extended and fully folded arm. There the columns are parallel, the parallelogram collapses to a segment, and $J$ loses rank. Those are the singular configurations — rank-deficient Jacobians, the subject the optimizer parts care about most.
Top: the joint-space square. Bottom: its sheared image at the current tip, with $J$'s columns as the edge arrows.
Quadratic forms: the shape in the eigenvalues
$q(x) = x^{\!\top} H x$ with $H$ symmetric
The second derivative of a scalar function of two variables is a $2\times2$ symmetric matrix, the Hessian $H$. The simplest such object is a quadratic form
Because $H$ is symmetric it has real eigenvalues $\lambda_1,\lambda_2$ and orthogonal eigenvectors. Those two numbers classify the shape completely: both positive is a bowl, both negative a cap, mixed signs a saddle, and a zero eigenvalue a flat valley or ridge (degenerate). Drag the entries; the contour map, the eigenvector axes and the 3D surface all follow, along with $\det H = ac - b^2$.
Contours of $q$ with the eigenvector axes drawn through the origin, and the surface $z=q(x,y)$ below.
The Hessian of a real function
The derivative of the gradient
Let $f(x,y)$ be any smooth scalar function. Its gradient $\nabla f$ is a vector map, so differentiating it gives a Jacobian again — and that matrix is the Hessian:
Drag the point and choose a function. The canvas shows the contour map, the gradient arrow (steepest ascent) and the eigenvalues of the numerically-computed Hessian. At a critical point — where $\nabla f = 0$ — the eigenvalues say everything: all positive is a local minimum, all negative a local maximum, mixed signs a saddle. Away from a critical point the same matrix measures the shape of the landscape underfoot.
Contours of $f$ with a draggable probe. The dark arrow is $\nabla f$; the readout gives $f$, $\nabla f$ and $H$ at the probe.
Why symmetry is the whole story
Mixed partials agree, so the eigenvalues are real
The Hessian is symmetric because the mixed partials commute: $f_{xy} = f_{yx}$ (Schwarz/Clairaut). Numerically the two orders are computed by different stencils — first in $x$ then $y$, versus first in $y$ then $x$ — and they agree to round-off. Drag the probe and compare them.
That symmetry is why the Hessian is not just any matrix. A real symmetric matrix has real eigenvalues and an orthogonal set of eigenvectors, so a surface really does have two well-defined principal curvatures, not complex ones. The two coloured L-shaped paths trace the same displacement in either order — and meet at the same corner. Part 14 builds the local quadratic model from exactly this matrix.
Two stencils for $f_{xy}$ and $f_{yx}$ drawn as paths from the probe: across then up (blue), up then across (pink). Same corner, same mixed partial.
Where this shows up
The matrix that runs the solver
Gauss–Newton — the engine behind Nonlinear Optimization, Part 1 — never inverts the residual map. It forms the Jacobian $J$ of residuals, solves the normal equations $J^{\!\top}J\,\delta = -J^{\!\top}r$, and steps. When the problem includes rotations (Part 2), unknown landmarks (Part 3) or loop closures (Part 4), the sparsity and rank of that same Jacobian decide whether the step is well posed. The Hessian's signs are the second-order check on top: a saddle has to be escaped, a bowl can be descended. Backpropagation in Volume II, Part 9 is the chain rule for these matrices, multiplied right-to-left.
Notation to carry forward
| Object | Reads as | Where it comes up |
|---|---|---|
J = ∂f/∂x | The matrix of first partials, $J_{ij}=\partial f_i/\partial x_j$ | Vector maps, Gauss–Newton, backprop |
f(x+δ) ≈ f(x) + Jδ | $J$ is the best linear approximation of a vector map | This part, Step 1 |
det J | Local area (or volume) scaling factor | Step 2; singular configurations at det J = 0 |
H = J(∇f) | The Hessian: derivative of the gradient, $H_{ij}=\partial^2 f/\partial x_i\partial x_j$ | Steps 3–4; Part 14 |
H = Hᵀ | Symmetric because $f_{xy}=f_{yx}$ | Step 5; guarantees real eigenvalues |
λ₁, λ₂ > 0 | Positive definite — local minimum (bowl) | The second-derivative test |
λ₁ λ₂ < 0 | Indefinite — saddle | Escaping saddle points in training |
Further reading
- 3Blue1Brown, The Jacobian matrix — the same columns as arrows, with the determinant as area.
- MIT 18.02, Multivariable Calculus — the course that proves the mixed partials commute.
- Wikipedia, Hessian matrix — the second-derivative test in full generality.