Zoom in until it's a line
The whole subject in one idea: almost every function is locally linear, and a derivative is just the slope you find when you zoom in far enough. Our running example is a robot moving along a straight line, with position $s(t) = \tfrac{1}{3}t^3 - \tfrac{3}{2}t^2 + 2t$ metres at time $t$ seconds. Everything below is one function, viewed from closer and closer.
Zoom in and it straightens out
The whole idea, in one slider
Look at the robot's position curve from far away and it bends. Put a magnifying glass on a point $a$ and keep turning it up: the piece of curve inside the glass looks flatter and flatter. What it flattens toward is the tangent line at $a$ — the straight line through $(a, s(a))$ whose slope is $s'(a)$. That is what "smooth" buys you, and it is the one fact the rest of the guide spends fourteen parts using.
Move $a$ and turn up the zoom. The window half-width $w$ shrinks by a factor of ten each step. The number to watch is not the error itself but the error per unit of window: if the curve is locally linear, the largest gap between curve and tangent shrinks like $w^2$, so dividing by $w^2$ leaves a bounded number — roughly $\tfrac{1}{2}|s''(a)|$.
The window $[a-w,\,a+w]$, rescaled to fill the frame. The curve (dark) and the tangent at a (teal) are drawn over it. As $w$ shrinks they merge.
A meter for linearity
How well can one line stand in for the curve?
Local linearity is a claim about a whole window, not just the slope at its centre. Take every point of the curve inside $[a-w, a+w]$, fit the single straight line that misses them least (least squares), and measure the worst miss as a fraction of the window's width:
Drag the probe anywhere along the curve, then shrink the window. On a smooth curve the misfit collapses to zero — the line fits as tightly as you like. That collapse is exactly what "locally linear" means, and it is why a tangent is a good approximation rather than merely a touching one.
Drag the probe along the curve. The shaded band is the window; the dashed line is the best fit inside it.
When zooming never straightens
Corners and oscillations
Smoothness is a real hypothesis, not decoration. Switch the function and keep zooming. The corner $|t-1.6|$ looks like a sharp $V$ at every scale: its best line keeps missing by a fixed fraction of the window, because the two sides have different slopes. The oscillation $\sin\!\bigl(1/(t-1.6)\bigr)$ is worse — it stays bounded in $[-1,1]$, yet no zoom tames it, because the wiggles only get faster as you approach the centre.
Both are continuous at $t=1.6$. Neither has a derivative there. The failure is precisely a failure of local linearity.
The same window around $t=1.6$, rescaled, for each function. Watch the misfit instead of the picture when the wiggle aliases.
The slope you find is the derivative
Why this is the whole subject
When the window shrank, the secant line through $a$ and $a+h$ rotated onto the tangent. The slope it settled on is a limit — the difference quotient's limit — and that limit is the derivative, $s'(a)$:
So "zoom in until it's a line" and "take the derivative" are two readings of one operation: the first is geometric, the second is arithmetic, and they agree to the last digit. Slide $a$ and $h$ and compare the secant slope, the tangent slope, and the closed form $s'(t)=t^2-3t+2$. The full development of that limit is Part 3; here it is enough that you have already seen it appear as the slope of the zoomed-in line.
The secant through a and a+h (dashed) rotates onto the tangent at a (solid) as h shrinks.
Where this shows up
The shape of the rest of the site
Every gradient-descent step is this page's tangent line. An optimizer cannot minimise a complicated function directly, so at the current point it builds the local line, steps downhill along it, and rebuilds. The whole method rests on the fact that near a point the function is its line — that is Nonlinear Optimization, and it assumes exactly what you just watched happen. The same fact, applied one composition at a time to a network, is backpropagation; the same fact, in many variables, is the Jacobian. Local linearity is the load-bearing assumption of the entire site.
What to carry forward
| Statement | Reads as | Where it comes up |
|---|---|---|
s(a+h) ≈ s(a) + s'(a)h | The local line: replace the curve by its tangent near a | Every linearisation, Newton, gradient descent |
L(t) = s(a) + s'(a)(t − a) | The tangent line through (a, s(a)) | The best straight-line stand-in |
|s(a+h) − L(a+h)| = O(h²) | Error shrinks one power faster than the step | Why one step of a solver is accurate |
s(a+h) = s(a) + s'(a)h + ½s''(a)h² + … | The curvature term that limits the line | Taylor, Part 7; curvature in optimisation |
f ∈ C¹ | Differentiable with continuous derivative — smooth enough to be locally linear | The regularity every solver quietly assumes |
|t − 1.6| | Continuous at 1.6 but never locally linear: a corner | The canonical non-example |
Further reading
- 3Blue1Brown, The paradox of the derivative — the same zoom, animated.
- Wikipedia, Linear approximation — the tangent-line statement and its error bound.
- MIT 18.01, Single Variable Calculus — the rigorous treatment that runs alongside this volume.