Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

The Jacobian of the arm

A vector map differentiated is a matrix

Our running example is a two-link planar arm. Its end-effector position is a vector-valued function of the joint angles,

$$ p(\theta_1,\theta_2) \;=\; \begin{bmatrix} L_1\cos\theta_1 + L_2\cos(\theta_1+\theta_2) \\[2pt] L_1\sin\theta_1 + L_2\sin(\theta_1+\theta_2) \end{bmatrix}. $$

Each output component gets its own derivative with respect to each input. Stacking them gives the Jacobian, the best linear approximation of $p$ near the current configuration:

$$ p(\theta+\delta) \;\approx\; p(\theta) + J\,\delta, \qquad J \;=\; \begin{bmatrix} \dfrac{\partial p_x}{\partial\theta_1} & \dfrac{\partial p_x}{\partial\theta_2} \\[8pt] \dfrac{\partial p_y}{\partial\theta_1} & \dfrac{\partial p_y}{\partial\theta_2} \end{bmatrix}. $$

This is Part 12's gradient, done once per output. The columns are especially readable: they are the velocity of the tip when only one joint moves. Drag the angles and watch the two arrows at the end-effector — they are those columns. Then set joint rates $\theta'$ and read the tip velocity $p' = J\theta'$.

Arm geometry with the two Jacobian columns drawn from the tip (blue = $\partial p/\partial\theta_1$, pink = $\partial p/\partial\theta_2$) and the actual tip velocity $p'=J\theta'$ in dark teal.

2

What J does to a neighbourhood

A square becomes a parallelogram; area scales by det J

Because $J$ is the local linear model, a tiny square of joint perturbations maps to a parallelogram of tip displacements. Left panel: a square of half-width $h$ in the $(\theta_1,\theta_2)$ plane. Right panel: its image under $J$, a sheared parallelogram — the columns of $J$ are its two edge directions.

The determinant is the area-scaling factor. Differentiating the arm formula gives the clean identity

$$ \det J \;=\; L_1 L_2 \sin\theta_2, $$

which vanishes at $\theta_2 = 0$ or $\theta_2 = \pi$: the fully extended and fully folded arm. There the columns are parallel, the parallelogram collapses to a segment, and $J$ loses rank. Those are the singular configurations — rank-deficient Jacobians, the subject the optimizer parts care about most.

Top: the joint-space square. Bottom: its sheared image at the current tip, with $J$'s columns as the edge arrows.

3

Quadratic forms: the shape in the eigenvalues

$q(x) = x^{\!\top} H x$ with $H$ symmetric

The second derivative of a scalar function of two variables is a $2\times2$ symmetric matrix, the Hessian $H$. The simplest such object is a quadratic form

$$ q(x,y) \;=\; \begin{bmatrix} x & y \end{bmatrix} \begin{bmatrix} a & b \\ b & c \end{bmatrix} \begin{bmatrix} x \\ y \end{bmatrix} \;=\; a\,x^2 + 2b\,xy + c\,y^2. $$

Because $H$ is symmetric it has real eigenvalues $\lambda_1,\lambda_2$ and orthogonal eigenvectors. Those two numbers classify the shape completely: both positive is a bowl, both negative a cap, mixed signs a saddle, and a zero eigenvalue a flat valley or ridge (degenerate). Drag the entries; the contour map, the eigenvector axes and the 3D surface all follow, along with $\det H = ac - b^2$.

Contours of $q$ with the eigenvector axes drawn through the origin, and the surface $z=q(x,y)$ below.

4

The Hessian of a real function

The derivative of the gradient

Let $f(x,y)$ be any smooth scalar function. Its gradient $\nabla f$ is a vector map, so differentiating it gives a Jacobian again — and that matrix is the Hessian:

$$ H(x) \;=\; J(\nabla f)(x) \;=\; \begin{bmatrix} f_{xx} & f_{xy} \\ f_{yx} & f_{yy} \end{bmatrix}. $$

Drag the point and choose a function. The canvas shows the contour map, the gradient arrow (steepest ascent) and the eigenvalues of the numerically-computed Hessian. At a critical point — where $\nabla f = 0$ — the eigenvalues say everything: all positive is a local minimum, all negative a local maximum, mixed signs a saddle. Away from a critical point the same matrix measures the shape of the landscape underfoot.

Contours of $f$ with a draggable probe. The dark arrow is $\nabla f$; the readout gives $f$, $\nabla f$ and $H$ at the probe.

5

Why symmetry is the whole story

Mixed partials agree, so the eigenvalues are real

The Hessian is symmetric because the mixed partials commute: $f_{xy} = f_{yx}$ (Schwarz/Clairaut). Numerically the two orders are computed by different stencils — first in $x$ then $y$, versus first in $y$ then $x$ — and they agree to round-off. Drag the probe and compare them.

That symmetry is why the Hessian is not just any matrix. A real symmetric matrix has real eigenvalues and an orthogonal set of eigenvectors, so a surface really does have two well-defined principal curvatures, not complex ones. The two coloured L-shaped paths trace the same displacement in either order — and meet at the same corner. Part 14 builds the local quadratic model from exactly this matrix.

Two stencils for $f_{xy}$ and $f_{yx}$ drawn as paths from the probe: across then up (blue), up then across (pink). Same corner, same mixed partial.

6

Where this shows up

The matrix that runs the solver

Gauss–Newton — the engine behind Nonlinear Optimization, Part 1 — never inverts the residual map. It forms the Jacobian $J$ of residuals, solves the normal equations $J^{\!\top}J\,\delta = -J^{\!\top}r$, and steps. When the problem includes rotations (Part 2), unknown landmarks (Part 3) or loop closures (Part 4), the sparsity and rank of that same Jacobian decide whether the step is well posed. The Hessian's signs are the second-order check on top: a saddle has to be escaped, a bowl can be descended. Backpropagation in Volume II, Part 9 is the chain rule for these matrices, multiplied right-to-left.

7

Notation to carry forward

ObjectReads asWhere it comes up
J = ∂f/∂xThe matrix of first partials, $J_{ij}=\partial f_i/\partial x_j$Vector maps, Gauss–Newton, backprop
f(x+δ) ≈ f(x) + Jδ$J$ is the best linear approximation of a vector mapThis part, Step 1
det JLocal area (or volume) scaling factorStep 2; singular configurations at det J = 0
H = J(∇f)The Hessian: derivative of the gradient, $H_{ij}=\partial^2 f/\partial x_i\partial x_j$Steps 3–4; Part 14
H = HᵀSymmetric because $f_{xy}=f_{yx}$Step 5; guarantees real eigenvalues
λ₁, λ₂ > 0Positive definite — local minimum (bowl)The second-derivative test
λ₁ λ₂ < 0Indefinite — saddleEscaping saddle points in training
8

Further reading

9

Check your understanding

0/4 answered