Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

The gradient collects the partials

Two slopes, one vector

For a function of two variables, $f(x,y)$, the partial derivatives $f_x$ and $f_y$ measure the slope you feel while walking due east and due north. Each is a number at each point, so each is a function. The gradient simply stacks them into a vector:

$$ \nabla f(x,y) \;=\; \bigl(f_x(x,y),\; f_y(x,y)\bigr). $$

Our running example is the two-link arm. Put the tip at $(x,y)$ and let the arm's objective be the squared distance from the tip to a target; written as a function of the two joint angles $(\theta_1, \theta_2)$ it is a landscape with hills and valleys, and its gradient says which way to nudge the joints to improve the objective fastest. For the picture it is easier to draw a height surface $f(x,y)$ directly: think of $f$ as the height of the ground. Everything below is the same mathematics either way.

The contour map is the ground viewed from above. Darker bands are higher ground; the thin curves are level sets, the places where $f$ is constant. Drag the point and the readout names its height and its gradient.

Contours of $f$ with the level curve through the point drawn solid. The pink arrow is the numeric gradient $\nabla f$, computed from $f$ alone by central differences.

2

The directional derivative

One number per direction, and it is a dot product

The partials answer only two questions: how fast does $f$ change going east, and going north? The useful question is how fast it changes in any direction. Let $\mathbf{u} = (u_1, u_2)$ be a unit vector, $|\mathbf{u}| = 1$, and walk from $\mathbf{p}$ along the straight line $\mathbf{r}(t) = \mathbf{p} + t\,\mathbf{u}$. The directional derivative is the rate of change of $f$ along that walk at the moment you set off:

$$ D_{\mathbf{u}} f(\mathbf{p}) \;=\; \frac{d}{dt}\, f(\mathbf{p} + t\,\mathbf{u})\Big|_{t=0}. $$

Apply the multivariable chain rule to the composition. The inner function has derivative $\mathbf{u}$, so the outer derivative — the pair of partials — contracts with it:

$$ D_{\mathbf{u}} f(\mathbf{p}) \;=\; f_x(\mathbf{p})\,u_1 + f_y(\mathbf{p})\,u_2 \;=\; \nabla f(\mathbf{p}) \cdot \mathbf{u}. $$

So the gradient is a machine that eats a direction and returns a rate: turn the dial to set $\mathbf{u}$, and the dot product is the slope in that direction. The panel below plots $D_{\mathbf{u}}f$ against the dial angle $\theta$ — it is a cosine, and everything about steepest ascent follows from that.

Top: the unit direction $\mathbf{u}$ (accent) from the draggable point, against the gradient (pink). Bottom: $D_{\mathbf{u}}f = |\nabla f|\cos\phi$ as the dial turns, with dashed lines at $\pm|\nabla f|$ and at zero.

3

Why the gradient is perpendicular to level sets

Differentiate the constraint $f(\mathbf{r}(t)) = c$

This is the conceptual centrepiece, and the proof is one line of the chain rule. A level set is a curve along which the value of $f$ never changes. Parametrise it as $\mathbf{r}(t)$ and write that fact down:

$$ f\bigl(\mathbf{r}(t)\bigr) \;=\; c \quad\text{for every } t. $$

Both sides are constant in $t$, so differentiate with respect to $t$. The chain rule turns the left side into a dot product with the tangent velocity $\mathbf{r}'(t)$, and the right side dies:

$$ \frac{d}{dt}\, f\bigl(\mathbf{r}(t)\bigr) \;=\; \nabla f\bigl(\mathbf{r}(t)\bigr)\cdot \mathbf{r}'(t) \;=\; 0. $$

That is the whole theorem: the gradient is orthogonal to the velocity of any path that stays on the level set. Since $\mathbf{r}'(t)$ is exactly the tangent direction, $\nabla f$ is normal to the level set. The demo makes it numeric — the tangent is estimated by finite differences along the curve, and their dot product hovers at zero wherever the point sits.

The point slides along one fixed level set. The gradient (pink) and the numerically-computed tangent (accent) meet at a right angle; the readout shows $\nabla f \cdot \mathbf{T} \approx 0$.

💡 The one-line proof, memorised: $f(\mathbf{r}(t)) = c \Rightarrow \nabla f \cdot \mathbf{r}' = 0$. A normal vector is precisely one that kills every tangent, so the gradient is normal to the level set — no geometry required.
4

Steepest ascent, and the direction of descent

Cauchy–Schwarz decides the best direction

Now combine the two facts. Since $D_{\mathbf{u}}f = \nabla f \cdot \mathbf{u}$ and $\mathbf{u}$ is a unit vector, Cauchy–Schwarz bounds the dot product by the product of lengths:

$$ D_{\mathbf{u}} f \;=\; |\nabla f|\,|\mathbf{u}|\cos\phi \;=\; |\nabla f|\cos\phi \;\le\; |\nabla f|, $$

where $\phi$ is the angle between $\mathbf{u}$ and $\nabla f$. Equality holds exactly when $\cos\phi = 1$, that is when $\mathbf{u}$ points the same way as $\nabla f$. So:

Feed that into an update rule and you get gradient descent, $\mathbf{p} \leftarrow \mathbf{p} - \eta\,\nabla f(\mathbf{p})$. Watch a particle take those steps both ways from one draggable start. Because every step is along $\nabla f$, the path crosses each contour at a right angle the whole way down (and up).

Pink runs uphill along $+\nabla f$ toward the peak; accent runs downhill along $-\nabla f$ toward the valley. Both start at the same draggable point.

⚠️ A fixed step size is a blunt instrument: too small and the particle crawls, too large and it overshoots the summit and rattles around it. Choosing $\eta$ well is the subject of Volume II, Part 10.
5

Where this shows up

This part justifies the optimizer

Every gradient-descent step on the site is Step 4 in disguise. The claim "move against the gradient to go downhill fastest" is not an approximation or a heuristic — it is the equality case of Cauchy–Schwarz, and it is why the very first method in Nonlinear Optimization takes the form $\mathbf{x} \leftarrow \mathbf{x} - \eta\,\nabla f(\mathbf{x})$. That page assumes the direction is right and asks how to choose the step; the reason the direction is right is here. How the curvature of $f$ makes some descents zig-zag and others crawl is Volume II, Part 10: Curvature and convergence. For the arm, the same arrow does real work: maximising reach toward a target is gradient ascent on a landscape whose coordinates are joint angles, so the vector that points uphill here points toward the pose that reaches.

6

Cheat sheet

ObjectStatementMeaning / use
Gradient$\nabla f = (f_x, f_y)$The stack of partials; a vector at every point
Directional derivative$D_{\mathbf{u}} f = \nabla f \cdot \mathbf{u}$Rate of change per unit step in direction $\mathbf{u}$, $|\mathbf{u}|=1$
Steepest ascent$\mathbf{u} = \nabla f / |\nabla f|$$D_{\mathbf{u}} f = +|\nabla f|$; equality in Cauchy–Schwarz
Steepest descent$\mathbf{u} = -\nabla f / |\nabla f|$$D_{\mathbf{u}} f = -|\nabla f|$; the optimizer's direction
Level-set tangency$\nabla f \cdot \mathbf{T} = 0$$\nabla f$ is normal to the level set; $\mathbf{T}$ is its tangent
Constraint differentiated$f(\mathbf{r}(t))=c \Rightarrow \nabla f \cdot \mathbf{r}' = 0$One chain-rule line proves the perpendicularity
Magnitude$|\nabla f|$The slope of the steepest path; zero at a critical point
7

Further reading

8

Check your understanding

0/4 answered