Linear approximation, Newton, Taylor
A tangent line is the first term of a much better idea: if a straight line is the best constant-slope stand-in for a curve, a quadratic is the best constant-curvature one, and a whole polynomial can be built to match the curve's derivatives at a point. Add terms and watch the match spread; then use the first-order version to run Newton's method and see its global behaviour turn fractal.
The linear model, measured
What the tangent buys you, and what it costs
Our running example is still the robot on a line, position $s(t) = \tfrac{1}{3}t^3 - \tfrac{3}{2}t^2 + 2t$ metres at time $t$ seconds. Part 3 ended on the observation that near a time $a$ the tangent line is the best straight-line stand-in for the curve. Make that quantitative. Move a step $h$ away from $a$ and the line predicts $s(a) + s'(a)h$; the curve delivers $s(a+h)$. The gap is the approximation error:
Because the first two terms cancel exactly, the error starts at order $h^2$: halve the step and the error drops by about four. Equivalently, $\text{error}/h^2$ settles on the constant $\tfrac{1}{2}s''(a)$. Slide $a$ and $h$ and watch that ratio hold even as the error itself vanishes.
Position against time. The teal line is the tangent at a; the pink bar at $a+h$ is the prediction error, curve minus line.
Add terms until the polynomial wraps the curve
Taylor's polynomial, derived by matching derivatives
The tangent matched two facts about $s$ at $a$: its value and its first derivative. Nothing stops us matching more. Ask for a polynomial $p(t) = c_0 + c_1(t-a) + c_2(t-a)^2 + c_3(t-a)^3 + \cdots$ whose derivatives at $a$ equal $s$'s. Evaluating at $t=a$ gives $c_0 = s(a)$. Differentiate once and set $t=a$: every higher term dies, leaving $c_1 = s'(a)$. Differentiate twice: the cubic term contributes $2c_2$ and everything above vanishes, so $c_2 = \tfrac{1}{2}s''(a)$. Differentiate $k$ times and only the $(t-a)^k$ term survives, contributing $k!\,c_k$:
That polynomial $p_n$ is the Taylor polynomial of order $n$ at $a$. It is the unique degree-$n$ polynomial that agrees with $s$ through its first $n$ derivatives there. The remainder quantifies what is left over; for some $\xi$ between $a$ and $a+h$ (Lagrange form),
So the error is controlled by the first derivative the polynomial did not match. Raise the order and the error shrinks — but only as far as the next derivative allows. The readout tracks the worst miss over a window around $a$.
The curve $s$ (dark) and its Taylor polynomial at a (teal, dashed). As the order rises the polynomial hugs the curve over a wider window.
Local versus global: the radius of convergence
More terms help — but only near $a$
A Taylor polynomial is built entirely from data at the single point $a$. How far it can be trusted is a global question, and the answer depends on the function. For a polynomial or for $\sin$, the series converges everywhere. For a function with a singularity, it stops at the distance from $a$ to the nearest one — even if the singularity is off the real line. The series for $1/(1+x^2)$ about $0$ is $1 - x^2 + x^4 - x^6 + \cdots$, and it converges only for $|x| < 1$: the poles at $x = \pm i$ sit at distance $1$. Inside the radius, more terms always help. Outside it, no number of terms does.
Pick a function and slide the center. The dashed red lines mark the radius $\sqrt{1+a^2}$ when it is finite; the readout reports the worst error over a window wide enough to cross it.
Curve (dark) and Taylor polynomial about a (teal dashed). Where the dashed curve flies away, the series has left its radius of convergence.
Newton's method: the linear model, set to zero
The first-order Taylor model as an algorithm
The linear model is not only for prediction; it is for solving. To find where $f(x) = 0$, stand at a guess $x_k$ and replace $f$ by its order-1 Taylor polynomial, $f(x_k) + f'(x_k)(x - x_k)$. That line is easy to zero. Setting it to zero and solving for the crossing point gives the next guess
This is Newton's method: draw the tangent, follow it to the axis, repeat. For the function with two roots $f(x) = x^2 - 1$ it is the ancient Babylonian iteration $x_{k+1} = \tfrac{1}{2}\bigl(x_k + 1/x_k\bigr)$, and the digits double roughly every step. Step it by hand, or press play; the iteration slider re-runs the whole sequence from the chosen start.
The curve $x^2-1$, the tangent at the current guess, and its intercept on the axis — the next guess.
Fast locally, fractal globally
Where Newton lands, as a picture
Near a simple root, Newton converges like a rocket — that is the local story of Step 4, and it is the same story an optimizer tells about its quadratic model. Zoom out, though, and the choice of starting point matters enormously. Colour every starting point by which root it reaches, allow the complex plane (the natural home of the roots of $z^n - 1$), and run the iteration to a fixed cap. The result is the basins of attraction: solid regions feeding each root, and boundaries that stay complicated at every scale.
Dark pixels are starts that escaped or had not converged within the cap. Slide the root count and the cap: more iterations reclassify the slow points, but the fractal frontier between basins never smooths out. Newton is a local method wearing global consequences.
Newton on $z^n - 1$, $z \in \mathbb{C}$, coloured by the root reached. The white horizontal line is the real axis.
Where this shows up
The optimizer's model, and the one-variable case
Every modern optimizer is a Taylor argument. Gradient descent keeps the linear term and walks downhill along it; Newton's method keeps the quadratic term, solves the resulting model exactly, and steps to its minimiser — which for the model is precisely $x - f'/f''$ in one variable. The multi-variable version is
and the page this whole thread leads to is Calculus in Motion: optimizers, where that model becomes a working step. For the descent and Newton recipes in their applied setting, see Nonlinear Optimization: its first part builds gradient descent and Newton from exactly the linear and quadratic models derived here. The matrix version of this expansion, and the second-derivative test that reads a landscape off the Hessian, is Part 14 — this page is its one-variable shadow.
Series worth memorising
| Function | Taylor series (about 0) | Converges for |
|---|---|---|
eˣ | Σ xⁿ / n! | every $x$ |
sin x | Σ (−1)ⁿ x^(2n+1) / (2n+1)! | every $x$ |
cos x | Σ (−1)ⁿ x^(2n) / (2n)! | every $x$ |
ln(1+x) | Σ (−1)^(n+1) xⁿ / n | $|x| < 1$, with $x = 1$ conditionally |
1/(1−x) | Σ xⁿ | $|x| < 1$ |
Further reading
- 3Blue1Brown, Taylor series — the polynomial-matching picture, animated through the same examples.
- Wikipedia, Taylor's theorem — the Lagrange remainder stated exactly as above.
- Wikipedia, Newton's method — convergence rates, and the first pictures of the fractal basins.