The gradient
Part 11 held one input fixed and let single-variable calculus back in as a partial derivative. Now let both inputs move at once. The partials stack into a single vector, the gradient, and that vector does two things worth a whole part: it points in the direction of steepest increase, and it stands perpendicular to every level set it touches. Drag a point across a contour map and watch both facts fall out.
The gradient collects the partials
Two slopes, one vector
For a function of two variables, $f(x,y)$, the partial derivatives $f_x$ and $f_y$ measure the slope you feel while walking due east and due north. Each is a number at each point, so each is a function. The gradient simply stacks them into a vector:
Our running example is the two-link arm. Put the tip at $(x,y)$ and let the arm's objective be the squared distance from the tip to a target; written as a function of the two joint angles $(\theta_1, \theta_2)$ it is a landscape with hills and valleys, and its gradient says which way to nudge the joints to improve the objective fastest. For the picture it is easier to draw a height surface $f(x,y)$ directly: think of $f$ as the height of the ground. Everything below is the same mathematics either way.
The contour map is the ground viewed from above. Darker bands are higher ground; the thin curves are level sets, the places where $f$ is constant. Drag the point and the readout names its height and its gradient.
Contours of $f$ with the level curve through the point drawn solid. The pink arrow is the numeric gradient $\nabla f$, computed from $f$ alone by central differences.
The directional derivative
One number per direction, and it is a dot product
The partials answer only two questions: how fast does $f$ change going east, and going north? The useful question is how fast it changes in any direction. Let $\mathbf{u} = (u_1, u_2)$ be a unit vector, $|\mathbf{u}| = 1$, and walk from $\mathbf{p}$ along the straight line $\mathbf{r}(t) = \mathbf{p} + t\,\mathbf{u}$. The directional derivative is the rate of change of $f$ along that walk at the moment you set off:
Apply the multivariable chain rule to the composition. The inner function has derivative $\mathbf{u}$, so the outer derivative — the pair of partials — contracts with it:
So the gradient is a machine that eats a direction and returns a rate: turn the dial to set $\mathbf{u}$, and the dot product is the slope in that direction. The panel below plots $D_{\mathbf{u}}f$ against the dial angle $\theta$ — it is a cosine, and everything about steepest ascent follows from that.
Top: the unit direction $\mathbf{u}$ (accent) from the draggable point, against the gradient (pink). Bottom: $D_{\mathbf{u}}f = |\nabla f|\cos\phi$ as the dial turns, with dashed lines at $\pm|\nabla f|$ and at zero.
Why the gradient is perpendicular to level sets
Differentiate the constraint $f(\mathbf{r}(t)) = c$
This is the conceptual centrepiece, and the proof is one line of the chain rule. A level set is a curve along which the value of $f$ never changes. Parametrise it as $\mathbf{r}(t)$ and write that fact down:
Both sides are constant in $t$, so differentiate with respect to $t$. The chain rule turns the left side into a dot product with the tangent velocity $\mathbf{r}'(t)$, and the right side dies:
That is the whole theorem: the gradient is orthogonal to the velocity of any path that stays on the level set. Since $\mathbf{r}'(t)$ is exactly the tangent direction, $\nabla f$ is normal to the level set. The demo makes it numeric — the tangent is estimated by finite differences along the curve, and their dot product hovers at zero wherever the point sits.
The point slides along one fixed level set. The gradient (pink) and the numerically-computed tangent (accent) meet at a right angle; the readout shows $\nabla f \cdot \mathbf{T} \approx 0$.
Steepest ascent, and the direction of descent
Cauchy–Schwarz decides the best direction
Now combine the two facts. Since $D_{\mathbf{u}}f = \nabla f \cdot \mathbf{u}$ and $\mathbf{u}$ is a unit vector, Cauchy–Schwarz bounds the dot product by the product of lengths:
where $\phi$ is the angle between $\mathbf{u}$ and $\nabla f$. Equality holds exactly when $\cos\phi = 1$, that is when $\mathbf{u}$ points the same way as $\nabla f$. So:
- the maximum is $+|\nabla f|$, at $\mathbf{u} = \nabla f/|\nabla f|$ — the gradient itself points uphill;
- the minimum is $-|\nabla f|$, at $\mathbf{u} = -\nabla f/|\nabla f|$ — so $-\nabla f$ is the steepest-descent direction;
- the value is zero when $\cos\phi = 0$, i.e. when $\mathbf{u}$ is tangent to the level set — walking along a contour neither climbs nor falls, which is Step 3 again.
Feed that into an update rule and you get gradient descent, $\mathbf{p} \leftarrow \mathbf{p} - \eta\,\nabla f(\mathbf{p})$. Watch a particle take those steps both ways from one draggable start. Because every step is along $\nabla f$, the path crosses each contour at a right angle the whole way down (and up).
Pink runs uphill along $+\nabla f$ toward the peak; accent runs downhill along $-\nabla f$ toward the valley. Both start at the same draggable point.
Where this shows up
This part justifies the optimizer
Every gradient-descent step on the site is Step 4 in disguise. The claim "move against the gradient to go downhill fastest" is not an approximation or a heuristic — it is the equality case of Cauchy–Schwarz, and it is why the very first method in Nonlinear Optimization takes the form $\mathbf{x} \leftarrow \mathbf{x} - \eta\,\nabla f(\mathbf{x})$. That page assumes the direction is right and asks how to choose the step; the reason the direction is right is here. How the curvature of $f$ makes some descents zig-zag and others crawl is Volume II, Part 10: Curvature and convergence. For the arm, the same arrow does real work: maximising reach toward a target is gradient ascent on a landscape whose coordinates are joint angles, so the vector that points uphill here points toward the pose that reaches.
Cheat sheet
| Object | Statement | Meaning / use |
|---|---|---|
| Gradient | $\nabla f = (f_x, f_y)$ | The stack of partials; a vector at every point |
| Directional derivative | $D_{\mathbf{u}} f = \nabla f \cdot \mathbf{u}$ | Rate of change per unit step in direction $\mathbf{u}$, $|\mathbf{u}|=1$ |
| Steepest ascent | $\mathbf{u} = \nabla f / |\nabla f|$ | $D_{\mathbf{u}} f = +|\nabla f|$; equality in Cauchy–Schwarz |
| Steepest descent | $\mathbf{u} = -\nabla f / |\nabla f|$ | $D_{\mathbf{u}} f = -|\nabla f|$; the optimizer's direction |
| Level-set tangency | $\nabla f \cdot \mathbf{T} = 0$ | $\nabla f$ is normal to the level set; $\mathbf{T}$ is its tangent |
| Constraint differentiated | $f(\mathbf{r}(t))=c \Rightarrow \nabla f \cdot \mathbf{r}' = 0$ | One chain-rule line proves the perpendicularity |
| Magnitude | $|\nabla f|$ | The slope of the steepest path; zero at a critical point |
Further reading
- MIT 18.02, Multivariable Calculus — the gradient, directional derivatives and the normal-to-level-set theorem, proved carefully.
- Khan Academy, Gradient and directional derivatives — extra practice with the dot-product reading.
- Wikipedia, Gradient — the reference form of every identity on this page.