The derivative
A derivative is a slope, and a slope is easy: rise over run. The whole idea is that as the run shrinks to nothing, the rise-over-run settles on a value. Drag the gap shut and watch it happen; then let the point slide along the curve and watch the derivative trace itself out underneath.
Average rate of change
Rise over run, over an interval
Our running example is a robot moving along a straight line. Its position at time $t$ is $s(t) = \tfrac{1}{3}t^3 - \tfrac{3}{2}t^2 + 2t$, in metres, for $t$ in seconds. Pick two times $a$ and $b$ and the average velocity over that window is the slope of the line joining the two points on the graph:
That line is called a secant line. It tells you what constant speed would cover the same ground in the same time. Drag either endpoint; the readout names the interval, the ground covered, and the average speed over it.
Position against time. The dashed line is the secant through a and b; the triangle shows the run and the rise.
Shrink the interval
The difference quotient, and its limit
Hold the first time $a$ fixed and slide the second time toward it. In the language of the interval, write $b = a + h$ so that $h$ is the width of the window. The average velocity becomes the difference quotient
and the derivative at $a$ is what this settles on as $h$ goes to zero:
That is the entire definition — no new machinery beyond a limit from Part 2. Shrinking $h$ slides the secant line around the fixed point until it lines up with the tangent. Watch the secant slope and the tangent slope in the readout converge.
The secant (dashed) through a and a+h, and the tangent (solid) at a. As h shrinks the dashed line rotates onto the solid one.
The derivative is itself a function
Slope at every point, assembled
The definition above gives a number for each choice of $a$. Let $a$ vary and those numbers assemble into a new function, $s'(t)$. The top panel is position; the bottom panel is the numerical derivative, sampled point by point with the same difference quotient. Slide $t$ and watch the two markers move together — the bottom marker's height is the slope of the top tangent.
Read the sign of the bottom curve as the robot's direction: positive means moving forward, negative means backing up, and the two places it crosses zero are where the robot is momentarily at rest.
Top: position $s(t)$ with the tangent at the current time. Bottom: its derivative $s'(t)$, computed numerically from the top curve alone.
When there is no derivative
Corners, cusps and jumps
The limit has to be one number. If the secant slope approaches one value from the left of $a$ and a different value from the right, the limit does not exist and the function is not differentiable there. A corner — the point of $|t-1.6|$ — is the standard example: slope $-1$ on one side, $+1$ on the other. A jump is worse: the difference quotient grows without bound. Switch the function and shrink $h$; the left and right slopes come from secant lines just either side of the probe.
The probe sits at t = 1.6. Left and right secants are drawn from a point h away on each side; where they disagree in the limit, there is no derivative.
The derivative is the best linear approximation
What the derivative is for
The slope is one reading of the derivative. The other, the one every later part leans on, is this: near $a$, the tangent line is the best straight-line stand-in for the curve, and the error shrinks like $h^2$ — one power faster than the distance you moved.
Move $h$ and watch the quadratic error. The ratio $\text{error}/h^2$ stays close to $\tfrac{1}{2}s''(a)$ — a constant. That is the local-linearity idea from Part 1 made quantitative, and it becomes the Taylor series of Part 7.
The curve, the tangent at a, and the prediction error at a+h drawn as a vertical bar.
Where this shows up
The shape of the rest of the site
Every gradient descent step is this part's tangent line: the optimizer replaces a function it cannot solve with the straight line it can, and follows the slope downhill. That is exactly Nonlinear Optimization, whose first parts assume this definition is comfortable. In machine learning the same limit, applied one composition at a time, is backpropagation — Volume II, Part 9 — and it is why training a network is just calculus done in the right order.
Notation to carry forward
| Notation | Reads as | Where it comes up |
|---|---|---|
f'(a), s'(t) | The derivative as a function, evaluated at a point | This volume; rates of change |
df/dx, ds/dt | Leibniz notation — a limit of a ratio, and the one that carries units | Related rates, the chain rule |
Δs / Δt | An average rate over a finite interval (the secant slope) | This part, Step 1 |
(f(a+h) − f(a)) / h | The difference quotient whose limit defines the derivative | The chain rule, Taylor, numerics |
f ∈ C¹ | f is differentiable and its derivative is continuous | The regularity every solver quietly assumes |
Further reading
- 3Blue1Brown, The paradox of the derivative — the visual argument for the same limit.
- MIT 18.01, Single Variable Calculus — the rigorous course that runs alongside this volume.
- Khan Academy, Differential calculus — extra drills on the difference quotient.