Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

The surface, and the first slice

A function $f:\mathbb{R}^2 \to \mathbb{R}$ is a landscape

A function of two variables takes a pair $(x,y)$ and returns a single height $z = f(x,y)$. Plot the height above the plane and you get a surface. Everything in Parts 3–10 happened on a curve: one input, one output, one slope. Here there are two input directions, so there are two questions to ask, and we answer each one by refusing to ask the other.

Pick a height $x = x_0$ and cut the landscape with a vertical plane along that line. On the cut, the surface becomes a plain curve in the $y$–$z$ plane — a function of the single variable $y$. Its ordinary slope at a point is the partial derivative with respect to $y$:

$$ \frac{\partial f}{\partial y}(x_0,y_0) \;=\; \lim_{h\to 0}\frac{f(x_0,\,y_0+h) - f(x_0,\,y_0)}{h}. $$

That is the whole definition: the single-variable derivative from Part 3, applied to the sliced curve while $x$ is held fixed and treated as a constant. Slide the plane and the sampled point; the readout shows $f$, the numeric $\partial f/\partial y$, and the exact value for our test surface $f(x,y)=x^2y+\sin y$.

Plotly surface $z=f(x,y)$. The translucent sheet is the slice plane $x=x_0$; the thick curve on it is the sliced function of $y$; the dot is the point being differentiated.

💡 Hold a variable fixed. The partial derivative is not a new kind of limit — it is exactly the old one, computed along the slice. That is why the rules of Parts 4–5 apply to it unchanged.
2

The other slice, and the total differential

Two slices, two partials, one linear model

Now cut with the other plane, $y = y_0$, and let the sliced curve run along $x$. Its slope is the partial with respect to $x$,

$$ \frac{\partial f}{\partial x}(x_0,y_0) \;=\; \lim_{h\to 0}\frac{f(x_0+h,\,y_0) - f(x_0,\,y_0)}{h}. $$

Neither partial alone can move the height in general; $f$ changes for a step that has both an $x$-part and a $y$-part. The single-variable linear approximation of Part 3 generalises to the total differential,

$$ df \;=\; \frac{\partial f}{\partial x}\,dx \;+\; \frac{\partial f}{\partial y}\,dy, $$

meaning $f(x_0+dx,\,y_0+dy) \approx f(x_0,y_0) + f_x\,dx + f_y\,dy$ up to error of second order in the step. Choose the slice, the probe, and two small steps $dx,dy$ and compare the true change with the linear prediction.

The slice plane $y=y_0$ and its curve along $x$. The readout compares the actual step with $f_x\,dx+f_y\,dy$.

3

The running example: a two-link arm

Forward kinematics is a function of two angles

From here on the running example is a planar robot arm with two links of lengths $L_1$ and $L_2$, joint angles $\theta_1$ (shoulder) and $\theta_2$ (elbow, relative to the first link). Its end-effector position is the tip you get by walking along the first link and then the second:

$$ \begin{aligned} x(\theta_1,\theta_2) &= L_1\cos\theta_1 + L_2\cos(\theta_1+\theta_2),\\ y(\theta_1,\theta_2) &= L_1\sin\theta_1 + L_2\sin(\theta_1+\theta_2). \end{aligned} $$

These are two functions of two variables, bundled into one vector output. Each partial answers a practical question: if I nudge only the elbow motor, how fast does the tip move in $x$? in $y$? The four numbers are computed numerically from the position functions alone (a finite difference in each angle) and shown beside the closed form.

The arm drawn in the plane. Change a link length or a joint angle; the tip moves and the four partials update in the readout.

4

A slice is just an ordinary curve

Parts 3–5 apply verbatim

The reason partials are easy is that there is no new rule. Freeze $x=x_0$ and define the one-variable function $g(y) = f(x_0,y)$ — the slice, lifted off the surface and laid flat. Then $\partial f/\partial y(x_0,y_0) = g'(y_0)$, and sums, products, the chain rule and every derivative from the toolbox apply to $g$ exactly as before. Change the slice and the whole curve changes, but along it you are back in single-variable calculus.

The top panel plots $g$; the bottom panel plots its derivative, sampled numerically from $g$ alone. The readout confirms the sliced derivative agrees with $\partial f/\partial y$ and with the closed form.

Top: the slice $g(y)=f(x_0,y)$ with its tangent at $y_0$. Bottom: $g'(y)$ computed from the top curve by finite differences — no $x$ dependence left in it.

5

Mixed partials don't care about order

A symmetry you can check numerically

Differentiate with respect to $x$, then with respect to $y$, and you get a second-order mixed partial $\partial^2 f/\partial y\,\partial x$. Do it in the other order and you get $\partial^2 f/\partial x\,\partial y$. For any function whose second partials are continuous — the $C^2$ case — the two are equal (Clairaut's theorem). This is why the Hessian of Part 13 is symmetric and why curvature has no preferred order.

Switch functions and slide the point. The two mixed partials are read off the finite-difference stencil in CalcViz.numericHess and reported separately; they land on top of each other to within round-off.

Both mixed partials sampled along $y$ at the current $x_0$. Solid is $\partial^2 f/\partial y\,\partial x$, dashed is $\partial^2 f/\partial x\,\partial y$; they coincide.

⚠️ The theorem needs continuity. If the second partials jump, the two orders can disagree; the equality is a regularity assumption, not a definition.
6

Regularity, and the reachable set

What “$C^1$” buys you

We will lean on one assumption constantly: that $f$ is $C^1$ — its partial derivatives exist and are continuous. Continuity of the partials is what makes the linear approximation honest,

$$ f(\mathbf{x}+\mathbf{h}) = f(\mathbf{x}) + \nabla f(\mathbf{x})\cdot\mathbf{h} + o(\lVert\mathbf{h}\rVert), $$

where $\nabla f = (f_x, f_y)$ is the gradient, the vector that collects the partials. The dot product is the preview of Part 12: the gradient is the single multivariable object that plays the role the derivative played in one dimension.

For the arm, $C^1$ is what makes the reachable set a surface. Sweep $\theta_1$ and $\theta_2$ across their ranges and the map $(\theta_1,\theta_2)\mapsto(x,y)$ paints a region of the plane. Where the map’s derivative matrix — the Jacobian of Part 13 — is non-singular, the map is locally a smooth change of coordinates and the region is a genuine surface patch rather than a fold. The partials of this step are exactly the entries of that matrix.

7

Where this shows up

Partial derivatives are the working currency

The moment a model has more than one parameter, the derivative becomes a vector of partials and every algorithm is written in terms of them. The next part, The gradient, assembles $f_x$ and $f_y$ into one arrow and shows it points straight uphill — which is the entire basis of gradient descent. That is precisely the loop in Nonlinear Optimization, where a solver repeatedly takes the partials of a loss and steps against them. For the arm, the same numbers become the Jacobian used to convert joint velocities into tip velocities.

8

Notation to carry forward

NotationReads asWhere it comes up
∂f/∂x, f_x, ∂_x fDifferentiate along $x$ with every other variable held fixedThis part; every multivariable model
∂f/∂y, f_yThe same, holding $x$ fixedSlices; the arm's elbow sensitivity
df = f_x dx + f_y dyThe total differential — the first-order change from a step in both inputsLinearisation, error propagation
∇f = (f_x, f_y)The gradient, a vector of partialsPart 12; gradient descent
∂²f/∂x∂y, f_xyDifferentiate in $x$ then in $y$Hessian, Part 13
f ∈ C¹Partials exist and are continuousThe regularity every solver assumes
9

Further reading

10

Check your understanding

0/4 answered