Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

The linear model, measured

What the tangent buys you, and what it costs

Our running example is still the robot on a line, position $s(t) = \tfrac{1}{3}t^3 - \tfrac{3}{2}t^2 + 2t$ metres at time $t$ seconds. Part 3 ended on the observation that near a time $a$ the tangent line is the best straight-line stand-in for the curve. Make that quantitative. Move a step $h$ away from $a$ and the line predicts $s(a) + s'(a)h$; the curve delivers $s(a+h)$. The gap is the approximation error:

$$ s(a+h) \;=\; \underbrace{s(a) + s'(a)\,h}_{\text{tangent line}} \;+\; \tfrac{1}{2}s''(a)\,h^2 + O(h^3). $$

Because the first two terms cancel exactly, the error starts at order $h^2$: halve the step and the error drops by about four. Equivalently, $\text{error}/h^2$ settles on the constant $\tfrac{1}{2}s''(a)$. Slide $a$ and $h$ and watch that ratio hold even as the error itself vanishes.

Position against time. The teal line is the tangent at a; the pink bar at $a+h$ is the prediction error, curve minus line.

2

Add terms until the polynomial wraps the curve

Taylor's polynomial, derived by matching derivatives

The tangent matched two facts about $s$ at $a$: its value and its first derivative. Nothing stops us matching more. Ask for a polynomial $p(t) = c_0 + c_1(t-a) + c_2(t-a)^2 + c_3(t-a)^3 + \cdots$ whose derivatives at $a$ equal $s$'s. Evaluating at $t=a$ gives $c_0 = s(a)$. Differentiate once and set $t=a$: every higher term dies, leaving $c_1 = s'(a)$. Differentiate twice: the cubic term contributes $2c_2$ and everything above vanishes, so $c_2 = \tfrac{1}{2}s''(a)$. Differentiate $k$ times and only the $(t-a)^k$ term survives, contributing $k!\,c_k$:

$$ c_k \;=\; \frac{s^{(k)}(a)}{k!}, \qquad p_n(t) \;=\; \sum_{k=0}^{n} \frac{s^{(k)}(a)}{k!}\,(t-a)^k. $$

That polynomial $p_n$ is the Taylor polynomial of order $n$ at $a$. It is the unique degree-$n$ polynomial that agrees with $s$ through its first $n$ derivatives there. The remainder quantifies what is left over; for some $\xi$ between $a$ and $a+h$ (Lagrange form),

$$ s(a+h) \;=\; p_n(a+h) + R_n(h), \qquad R_n(h) \;=\; \frac{s^{(n+1)}(\xi)}{(n+1)!}\,h^{n+1}. $$

So the error is controlled by the first derivative the polynomial did not match. Raise the order and the error shrinks — but only as far as the next derivative allows. The readout tracks the worst miss over a window around $a$.

The curve $s$ (dark) and its Taylor polynomial at a (teal, dashed). As the order rises the polynomial hugs the curve over a wider window.

💡 Terminating series: the robot's position is a cubic, so once $n \ge 3$ the remainder is exactly zero and every extra term is $0\cdot h^k$. A Taylor polynomial can reproduce a polynomial exactly; for the transcendental functions in the next step it never stops.
3

Local versus global: the radius of convergence

More terms help — but only near $a$

A Taylor polynomial is built entirely from data at the single point $a$. How far it can be trusted is a global question, and the answer depends on the function. For a polynomial or for $\sin$, the series converges everywhere. For a function with a singularity, it stops at the distance from $a$ to the nearest one — even if the singularity is off the real line. The series for $1/(1+x^2)$ about $0$ is $1 - x^2 + x^4 - x^6 + \cdots$, and it converges only for $|x| < 1$: the poles at $x = \pm i$ sit at distance $1$. Inside the radius, more terms always help. Outside it, no number of terms does.

Pick a function and slide the center. The dashed red lines mark the radius $\sqrt{1+a^2}$ when it is finite; the readout reports the worst error over a window wide enough to cross it.

Curve (dark) and Taylor polynomial about a (teal dashed). Where the dashed curve flies away, the series has left its radius of convergence.

4

Newton's method: the linear model, set to zero

The first-order Taylor model as an algorithm

The linear model is not only for prediction; it is for solving. To find where $f(x) = 0$, stand at a guess $x_k$ and replace $f$ by its order-1 Taylor polynomial, $f(x_k) + f'(x_k)(x - x_k)$. That line is easy to zero. Setting it to zero and solving for the crossing point gives the next guess

$$ x_{k+1} \;=\; x_k - \frac{f(x_k)}{f'(x_k)}. $$

This is Newton's method: draw the tangent, follow it to the axis, repeat. For the function with two roots $f(x) = x^2 - 1$ it is the ancient Babylonian iteration $x_{k+1} = \tfrac{1}{2}\bigl(x_k + 1/x_k\bigr)$, and the digits double roughly every step. Step it by hand, or press play; the iteration slider re-runs the whole sequence from the chosen start.

The curve $x^2-1$, the tangent at the current guess, and its intercept on the axis — the next guess.

5

Fast locally, fractal globally

Where Newton lands, as a picture

Near a simple root, Newton converges like a rocket — that is the local story of Step 4, and it is the same story an optimizer tells about its quadratic model. Zoom out, though, and the choice of starting point matters enormously. Colour every starting point by which root it reaches, allow the complex plane (the natural home of the roots of $z^n - 1$), and run the iteration to a fixed cap. The result is the basins of attraction: solid regions feeding each root, and boundaries that stay complicated at every scale.

Dark pixels are starts that escaped or had not converged within the cap. Slide the root count and the cap: more iterations reclassify the slow points, but the fractal frontier between basins never smooths out. Newton is a local method wearing global consequences.

Newton on $z^n - 1$, $z \in \mathbb{C}$, coloured by the root reached. The white horizontal line is the real axis.

6

Where this shows up

The optimizer's model, and the one-variable case

Every modern optimizer is a Taylor argument. Gradient descent keeps the linear term and walks downhill along it; Newton's method keeps the quadratic term, solves the resulting model exactly, and steps to its minimiser — which for the model is precisely $x - f'/f''$ in one variable. The multi-variable version is

$$ f(x_k + p) \;\approx\; f(x_k) + \nabla f(x_k)^\top p + \tfrac{1}{2}\,p^\top \nabla^2 f(x_k)\,p, $$

and the page this whole thread leads to is Calculus in Motion: optimizers, where that model becomes a working step. For the descent and Newton recipes in their applied setting, see Nonlinear Optimization: its first part builds gradient descent and Newton from exactly the linear and quadratic models derived here. The matrix version of this expansion, and the second-derivative test that reads a landscape off the Hessian, is Part 14 — this page is its one-variable shadow.

7

Series worth memorising

FunctionTaylor series (about 0)Converges for
eˣΣ xⁿ / n!every $x$
sin xΣ (−1)ⁿ x^(2n+1) / (2n+1)!every $x$
cos xΣ (−1)ⁿ x^(2n) / (2n)!every $x$
ln(1+x)Σ (−1)^(n+1) xⁿ / n$|x| < 1$, with $x = 1$ conditionally
1/(1−x)Σ xⁿ$|x| < 1$
8

Further reading

9

Check your understanding

0/4 answered