Functions of several variables
Everything so far differentiated a function of one number. A robot arm does not have one number: give it two joint angles and its end-effector position is a function of two inputs at once. The trick that keeps single-variable calculus alive is to freeze all but one input. What is left is an ordinary curve, and its ordinary slope is a partial derivative. Do that for every input and you have the machinery for the rest of the volume.
The surface, and the first slice
A function $f:\mathbb{R}^2 \to \mathbb{R}$ is a landscape
A function of two variables takes a pair $(x,y)$ and returns a single height $z = f(x,y)$. Plot the height above the plane and you get a surface. Everything in Parts 3–10 happened on a curve: one input, one output, one slope. Here there are two input directions, so there are two questions to ask, and we answer each one by refusing to ask the other.
Pick a height $x = x_0$ and cut the landscape with a vertical plane along that line. On the cut, the surface becomes a plain curve in the $y$–$z$ plane — a function of the single variable $y$. Its ordinary slope at a point is the partial derivative with respect to $y$:
That is the whole definition: the single-variable derivative from Part 3, applied to the sliced curve while $x$ is held fixed and treated as a constant. Slide the plane and the sampled point; the readout shows $f$, the numeric $\partial f/\partial y$, and the exact value for our test surface $f(x,y)=x^2y+\sin y$.
Plotly surface $z=f(x,y)$. The translucent sheet is the slice plane $x=x_0$; the thick curve on it is the sliced function of $y$; the dot is the point being differentiated.
The other slice, and the total differential
Two slices, two partials, one linear model
Now cut with the other plane, $y = y_0$, and let the sliced curve run along $x$. Its slope is the partial with respect to $x$,
Neither partial alone can move the height in general; $f$ changes for a step that has both an $x$-part and a $y$-part. The single-variable linear approximation of Part 3 generalises to the total differential,
meaning $f(x_0+dx,\,y_0+dy) \approx f(x_0,y_0) + f_x\,dx + f_y\,dy$ up to error of second order in the step. Choose the slice, the probe, and two small steps $dx,dy$ and compare the true change with the linear prediction.
The slice plane $y=y_0$ and its curve along $x$. The readout compares the actual step with $f_x\,dx+f_y\,dy$.
The running example: a two-link arm
Forward kinematics is a function of two angles
From here on the running example is a planar robot arm with two links of lengths $L_1$ and $L_2$, joint angles $\theta_1$ (shoulder) and $\theta_2$ (elbow, relative to the first link). Its end-effector position is the tip you get by walking along the first link and then the second:
These are two functions of two variables, bundled into one vector output. Each partial answers a practical question: if I nudge only the elbow motor, how fast does the tip move in $x$? in $y$? The four numbers are computed numerically from the position functions alone (a finite difference in each angle) and shown beside the closed form.
The arm drawn in the plane. Change a link length or a joint angle; the tip moves and the four partials update in the readout.
A slice is just an ordinary curve
Parts 3–5 apply verbatim
The reason partials are easy is that there is no new rule. Freeze $x=x_0$ and define the one-variable function $g(y) = f(x_0,y)$ — the slice, lifted off the surface and laid flat. Then $\partial f/\partial y(x_0,y_0) = g'(y_0)$, and sums, products, the chain rule and every derivative from the toolbox apply to $g$ exactly as before. Change the slice and the whole curve changes, but along it you are back in single-variable calculus.
The top panel plots $g$; the bottom panel plots its derivative, sampled numerically from $g$ alone. The readout confirms the sliced derivative agrees with $\partial f/\partial y$ and with the closed form.
Top: the slice $g(y)=f(x_0,y)$ with its tangent at $y_0$. Bottom: $g'(y)$ computed from the top curve by finite differences — no $x$ dependence left in it.
Mixed partials don't care about order
A symmetry you can check numerically
Differentiate with respect to $x$, then with respect to $y$, and you get a second-order mixed partial $\partial^2 f/\partial y\,\partial x$. Do it in the other order and you get $\partial^2 f/\partial x\,\partial y$. For any function whose second partials are continuous — the $C^2$ case — the two are equal (Clairaut's theorem). This is why the Hessian of Part 13 is symmetric and why curvature has no preferred order.
Switch functions and slide the point. The two mixed partials are read off the finite-difference stencil in CalcViz.numericHess and reported separately; they land on top of each other to within round-off.
Both mixed partials sampled along $y$ at the current $x_0$. Solid is $\partial^2 f/\partial y\,\partial x$, dashed is $\partial^2 f/\partial x\,\partial y$; they coincide.
Regularity, and the reachable set
What “$C^1$” buys you
We will lean on one assumption constantly: that $f$ is $C^1$ — its partial derivatives exist and are continuous. Continuity of the partials is what makes the linear approximation honest,
where $\nabla f = (f_x, f_y)$ is the gradient, the vector that collects the partials. The dot product is the preview of Part 12: the gradient is the single multivariable object that plays the role the derivative played in one dimension.
For the arm, $C^1$ is what makes the reachable set a surface. Sweep $\theta_1$ and $\theta_2$ across their ranges and the map $(\theta_1,\theta_2)\mapsto(x,y)$ paints a region of the plane. Where the map’s derivative matrix — the Jacobian of Part 13 — is non-singular, the map is locally a smooth change of coordinates and the region is a genuine surface patch rather than a fold. The partials of this step are exactly the entries of that matrix.
Where this shows up
Partial derivatives are the working currency
The moment a model has more than one parameter, the derivative becomes a vector of partials and every algorithm is written in terms of them. The next part, The gradient, assembles $f_x$ and $f_y$ into one arrow and shows it points straight uphill — which is the entire basis of gradient descent. That is precisely the loop in Nonlinear Optimization, where a solver repeatedly takes the partials of a loss and steps against them. For the arm, the same numbers become the Jacobian used to convert joint velocities into tip velocities.
Notation to carry forward
| Notation | Reads as | Where it comes up |
|---|---|---|
∂f/∂x, f_x, ∂_x f | Differentiate along $x$ with every other variable held fixed | This part; every multivariable model |
∂f/∂y, f_y | The same, holding $x$ fixed | Slices; the arm's elbow sensitivity |
df = f_x dx + f_y dy | The total differential — the first-order change from a step in both inputs | Linearisation, error propagation |
∇f = (f_x, f_y) | The gradient, a vector of partials | Part 12; gradient descent |
∂²f/∂x∂y, f_xy | Differentiate in $x$ then in $y$ | Hessian, Part 13 |
f ∈ C¹ | Partials exist and are continuous | The regularity every solver assumes |
Further reading
- 3Blue1Brown, Multivariable calculus — the geometric picture of slices and the gradient.
- MIT 18.02, Multivariable Calculus — the rigorous companion to this volume.
- Khan Academy, Multivariable calculus — drills on partial derivatives and the chain rule in several variables.