Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

A curved space that is locally flat

The manifold, the chart, and the tangent space

Our running example has been the two-link planar arm reaching for a target. The direction the tool points is a unit vector, and the set of all unit vectors in three dimensions is the unit sphere $S^2$. That sphere is a manifold: a space that is not flat globally, but every point has an open neighbourhood that a chart identifies with an open piece of $\mathbb{R}^n$. Locally it is $\mathbb{R}^2$ — longitude and latitude are the chart. There is no single chart covering the whole sphere, which is exactly why differentiation has to be assembled point by point.

Fix a point $p$ on the sphere. Directions you may move while staying on the surface are the vectors orthogonal to the outward normal $p$. Those vectors form the tangent space, a genuine vector space attached at $p$:

$$ T_p M \;=\; \{\, v \in \mathbb{R}^3 : v \cdot p = 0 \,\}. $$

Now put a scalar field on the sphere, say $f(p) = 0.7\,p_z + 0.35\,p_x p_y$ (a height plus a saddle). Its ordinary gradient $\nabla f$ is a vector in the ambient $\mathbb{R}^3$, and it generally has a component poking out of the surface — a direction you cannot move. The derivative that lives on the manifold is the Riemannian gradient: the ambient gradient with its normal part removed,

$$ \nabla_M f(p) \;=\; \nabla f(p) - \bigl(\nabla f(p)\cdot p\bigr)\,p . $$

Drag the point $p$ around the sphere. The tangent plane, an orthonormal basis $e_1, e_2$ for it, the ambient gradient (purple) and the tangential gradient (amber) all follow, and the readout prints each vector plus how much of $\nabla f$ had to be thrown away.

Orthographic view of $S^2$. Drag the dark point: the tangent plane and basis rotate with it. Purple is the ambient $\nabla f$; amber is its projection $\nabla_M f$ onto the plane.

💡 Ambient versus Riemannian. The ambient gradient answers "which way is uphill in $\mathbb{R}^3$". The Riemannian gradient answers "which way is uphill while staying on the surface". Only the second is a legal search direction on the manifold.
2

Moving while staying on the surface

The exponential map and retractions

A tangent vector $v \in T_pM$ is a direction and a length, but $p + v$ is generally off the surface. Two ways to land back on $M$. The exponential map walks the straightest possible path — the geodesic — for arc length $|v|$:

$$ \exp_p(v) \;=\; \cos|v|\; p + \frac{\sin|v|}{|v|}\; v . $$

A retraction is any cheaper map that agrees with it to first order. The simplest is to add and renormalise:

$$ R_p(v) \;=\; \frac{p + v}{\lVert p + v \rVert}. $$

For the sphere the retraction is not random: $R_p(v)$ lands on the same geodesic as $\exp_p$, but at angle $\arctan|v|$ instead of $|v|$. So $R_p(v) = \exp_p\!\bigl(\tfrac{\arctan|v|}{|v|}\,v\bigr)$ — identical to first order, wrong at third order. Slide the step $h$ and the direction; the readout compares the two norms, the two geodesic distances from $p$, and the chord distance between the two landing points.

A tangent step $v$ from $p$. Blue lands by the exponential map, magenta by add-and-renormalise, the hollow dot is the illegal $p+v$ floating off the sphere.

⚠️ Any retraction is only first-order accurate. At small $h$ the two landing points look identical; the geodesic distances already differ by $h - \arctan h \approx h^3/3$.
3

Why rotations do not add

SO(2) adds; SO(3) does not

In the plane, composing two rotations adds their angles and order does not matter: $R(\alpha)R(\beta) = R(\alpha+\beta)$, because $SO(2)$ is commutative. That is the exception. In three dimensions, rotating about $x$ and then about $y$ is not the same as the other order, and neither is a rotation by $\alpha+\beta$:

$$ R_x(\alpha)\,R_y(\beta) \;\neq\; R_y(\beta)\,R_x(\alpha) \quad\text{in } SO(3). $$

The leftover is the commutator of the two generators — the first term of the Baker–Campbell–Hausdorff series — so rotation parameters cannot be added like numbers. The correct picture is always: compose in the group (multiply matrices), and perturb in the tangent space (the Lie algebra $\mathfrak{so}(3)$). We do not re-derive Rodrigues' formula or the matrix exponential here; the rotation primer linked at the end builds both from scratch. The canvas is the friendly $SO(2)$ case, where the two hands coincide exactly; the readout also tracks the $SO(3)$ analogue computed in the standard axis–angle representation.

In $SO(2)$ the composed hand (magenta) sits exactly on the dashed $\alpha+\beta$ hand. Toggle the order: nothing moves, because planar rotations commute — the readout shows where 3D stops agreeing.

4

Retractions drift

First-order accuracy is not enough over many steps

One step hides the difference; a loop exposes it. Say we want to follow a tangent field around the sphere — the flow that rotates about a fixed axis. Each step we evaluate the field, take a step of size $h$, and either exponentiate or renormalise. Both integrate the same ODE to first order, with local error $O(h^2)$ per step, so over $K$ steps the two trajectories separate by $O(Kh^2)$:

$$ \text{one step: } O(h^2), \qquad \text{over } K \text{ steps: } O(K h^2). $$

Press Play to walk both around the loop. The readout accumulates the geodesic separation step by step and reports how badly each trajectory fails to close. Enlarge the step and the two retractions part company faster — this is why production solvers use a retraction with the right geometry, not merely any projection.

Blue: exponential-map steps. Magenta: add-and-renormalise steps. Both chase the same axis-rotation field for one revolution.

5

Optimising on the manifold

Riemannian gradient descent

Now put the pieces together. To minimise $f$ on $M$: project the ambient gradient into the tangent space, step downhill there, and retract back onto the surface. That is Riemannian gradient descent, and it is the flat algorithm with one projection added:

$$ p_{k+1} \;=\; \exp_{p_k}\!\bigl(-\eta\,\nabla_M f(p_k)\bigr) \quad\text{or}\quad p_{k+1} \;=\; R_{p_k}\!\bigl(-\eta\,\nabla_M f(p_k)\bigr). $$

The two runs below minimise $f(p) = 0.7\,p_z + 0.35\,p_x p_y$ on the same sphere from the same start, one exponentiating and one renormalising. The exact minimum is the south pole $p^\star = (0,0,-1)$, where $\nabla f = -0.7\,p^\star$ is purely normal and the tangential gradient vanishes. Slide the learning rate: both get there, but the retraction path peels away from the geodesic one. This is the machinery behind fitting rotations — rotations are a manifold too, and its tangent-space perturbation is built in the rotation primer.

Blue: exponential-map descent. Magenta: renormalised descent. The amber ring marks the minimum at the south pole.

6

Where this shows up

From a sphere to a rotation

The sphere is the toy; rotations are the job. Optimising a camera pose, a robot joint, or a neural network on $SO(3)$ uses exactly the three moves above — a tangent space, an exponential map or retraction, and a Riemannian gradient. The matrix exponential, Rodrigues' formula, the tangent-space perturbation $R\cdot\exp(\delta)$ and the Jacobian that drives Gauss–Newton on a rotation are all worked out in SO(3) From Scratch, and the full theory, including poses, the adjoint and uncertainty, is the Lie Groups & Lie Algebras course. The hill-walking itself — steps, momentum, Newton and trust regions, now on a curved surface — is Nonlinear Optimization, and the flat theory of why those steps behave as they do was Curvature and convergence. On the arm, the joint vector is a chart for a torus, and every "reach" is a descent on that manifold.

7

Notation to carry forward

ObjectReads asWhere it comes up
M, TpMA manifold and its tangent space at p — the local Rn of directionsEvery curved parameter space
∇Mf = ∇f − (∇f·p)pRiemannian gradient: the ambient gradient minus its normal partThe only legal descent direction on M
expp(v)Geodesic from p with initial velocity v, length |v|Exact surface moves; rotations via the matrix exponential
Rp(v)A retraction — any first-order stand-in for exp, e.g. (p+v)/|p+v|Cheap updates when exp is expensive
geodesic vs chordArc distance acos(p·q) versus straight-line |p−q|Measuring error without leaving the surface
R(α)R(β) vs R(α+β)Compose versus add angles — equal in SO(2), unequal in SO(3)Why rotation parameters must not be added
pk+1 = exppk(−η∇Mf)Riemannian gradient descentLearning on rotations, poses and the sphere
8

Further reading

9

Check your understanding

0/4 answered