Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

Zoom in and it straightens out

The whole idea, in one slider

Look at the robot's position curve from far away and it bends. Put a magnifying glass on a point $a$ and keep turning it up: the piece of curve inside the glass looks flatter and flatter. What it flattens toward is the tangent line at $a$ — the straight line through $(a, s(a))$ whose slope is $s'(a)$. That is what "smooth" buys you, and it is the one fact the rest of the guide spends fourteen parts using.

Move $a$ and turn up the zoom. The window half-width $w$ shrinks by a factor of ten each step. The number to watch is not the error itself but the error per unit of window: if the curve is locally linear, the largest gap between curve and tangent shrinks like $w^2$, so dividing by $w^2$ leaves a bounded number — roughly $\tfrac{1}{2}|s''(a)|$.

$$ s(a+h) - \bigl[s(a) + s'(a)h\bigr] \;=\; \tfrac{1}{2}s''(a)\,h^2 + O(h^3). $$

The window $[a-w,\,a+w]$, rescaled to fill the frame. The curve (dark) and the tangent at a (teal) are drawn over it. As $w$ shrinks they merge.

2

A meter for linearity

How well can one line stand in for the curve?

Local linearity is a claim about a whole window, not just the slope at its centre. Take every point of the curve inside $[a-w, a+w]$, fit the single straight line that misses them least (least squares), and measure the worst miss as a fraction of the window's width:

$$ \text{misfit}(a,w) \;=\; \frac{\max_{|t-a|\le w}\bigl|\,s(t) - \text{best line}(t)\,\bigr|}{w}. $$

Drag the probe anywhere along the curve, then shrink the window. On a smooth curve the misfit collapses to zero — the line fits as tightly as you like. That collapse is exactly what "locally linear" means, and it is why a tangent is a good approximation rather than merely a touching one.

Drag the probe along the curve. The shaded band is the window; the dashed line is the best fit inside it.

💡 Read it as a limit: misfit is a number attached to every window width. Smooth means misfit $\to 0$ as $w \to 0$; a corner keeps a fixed misfit no matter how small the window gets. That is the entire difference between differentiable and not.
3

When zooming never straightens

Corners and oscillations

Smoothness is a real hypothesis, not decoration. Switch the function and keep zooming. The corner $|t-1.6|$ looks like a sharp $V$ at every scale: its best line keeps missing by a fixed fraction of the window, because the two sides have different slopes. The oscillation $\sin\!\bigl(1/(t-1.6)\bigr)$ is worse — it stays bounded in $[-1,1]$, yet no zoom tames it, because the wiggles only get faster as you approach the centre.

Both are continuous at $t=1.6$. Neither has a derivative there. The failure is precisely a failure of local linearity.

The same window around $t=1.6$, rescaled, for each function. Watch the misfit instead of the picture when the wiggle aliases.

⚠️ Zooming is a microscope, not a proof. It makes local linearity visible and believable; the epsilon-delta argument that pins it down is Part 2.
4

The slope you find is the derivative

Why this is the whole subject

When the window shrank, the secant line through $a$ and $a+h$ rotated onto the tangent. The slope it settled on is a limit — the difference quotient's limit — and that limit is the derivative, $s'(a)$:

$$ s'(a) \;=\; \lim_{h\to 0}\frac{s(a+h)-s(a)}{h}. $$

So "zoom in until it's a line" and "take the derivative" are two readings of one operation: the first is geometric, the second is arithmetic, and they agree to the last digit. Slide $a$ and $h$ and compare the secant slope, the tangent slope, and the closed form $s'(t)=t^2-3t+2$. The full development of that limit is Part 3; here it is enough that you have already seen it appear as the slope of the zoomed-in line.

The secant through a and a+h (dashed) rotates onto the tangent at a (solid) as h shrinks.

5

Where this shows up

The shape of the rest of the site

Every gradient-descent step is this page's tangent line. An optimizer cannot minimise a complicated function directly, so at the current point it builds the local line, steps downhill along it, and rebuilds. The whole method rests on the fact that near a point the function is its line — that is Nonlinear Optimization, and it assumes exactly what you just watched happen. The same fact, applied one composition at a time to a network, is backpropagation; the same fact, in many variables, is the Jacobian. Local linearity is the load-bearing assumption of the entire site.

6

What to carry forward

StatementReads asWhere it comes up
s(a+h) ≈ s(a) + s'(a)hThe local line: replace the curve by its tangent near aEvery linearisation, Newton, gradient descent
L(t) = s(a) + s'(a)(t − a)The tangent line through (a, s(a))The best straight-line stand-in
|s(a+h) − L(a+h)| = O(h²)Error shrinks one power faster than the stepWhy one step of a solver is accurate
s(a+h) = s(a) + s'(a)h + ½s''(a)h² + …The curvature term that limits the lineTaylor, Part 7; curvature in optimisation
f ∈ C¹Differentiable with continuous derivative — smooth enough to be locally linearThe regularity every solver quietly assumes
|t − 1.6|Continuous at 1.6 but never locally linear: a cornerThe canonical non-example
7

Further reading

8

Check your understanding

0/4 answered