Calculus on manifolds
Everything so far has lived in flat space, where you may add vectors and move in straight lines. A sphere, a rotation, the set of unit directions a tool can point along — these are curved, and none of them is a vector space. This page rebuilds the three pieces of calculus that survive: the tangent space, the gradient, and the step. The naive step and the exponential-map step agree to first order and disagree forever after, and that difference is what every optimizer on rotations has to get right.
A curved space that is locally flat
The manifold, the chart, and the tangent space
Our running example has been the two-link planar arm reaching for a target. The direction the tool points is a unit vector, and the set of all unit vectors in three dimensions is the unit sphere $S^2$. That sphere is a manifold: a space that is not flat globally, but every point has an open neighbourhood that a chart identifies with an open piece of $\mathbb{R}^n$. Locally it is $\mathbb{R}^2$ — longitude and latitude are the chart. There is no single chart covering the whole sphere, which is exactly why differentiation has to be assembled point by point.
Fix a point $p$ on the sphere. Directions you may move while staying on the surface are the vectors orthogonal to the outward normal $p$. Those vectors form the tangent space, a genuine vector space attached at $p$:
Now put a scalar field on the sphere, say $f(p) = 0.7\,p_z + 0.35\,p_x p_y$ (a height plus a saddle). Its ordinary gradient $\nabla f$ is a vector in the ambient $\mathbb{R}^3$, and it generally has a component poking out of the surface — a direction you cannot move. The derivative that lives on the manifold is the Riemannian gradient: the ambient gradient with its normal part removed,
Drag the point $p$ around the sphere. The tangent plane, an orthonormal basis $e_1, e_2$ for it, the ambient gradient (purple) and the tangential gradient (amber) all follow, and the readout prints each vector plus how much of $\nabla f$ had to be thrown away.
Orthographic view of $S^2$. Drag the dark point: the tangent plane and basis rotate with it. Purple is the ambient $\nabla f$; amber is its projection $\nabla_M f$ onto the plane.
Moving while staying on the surface
The exponential map and retractions
A tangent vector $v \in T_pM$ is a direction and a length, but $p + v$ is generally off the surface. Two ways to land back on $M$. The exponential map walks the straightest possible path — the geodesic — for arc length $|v|$:
A retraction is any cheaper map that agrees with it to first order. The simplest is to add and renormalise:
For the sphere the retraction is not random: $R_p(v)$ lands on the same geodesic as $\exp_p$, but at angle $\arctan|v|$ instead of $|v|$. So $R_p(v) = \exp_p\!\bigl(\tfrac{\arctan|v|}{|v|}\,v\bigr)$ — identical to first order, wrong at third order. Slide the step $h$ and the direction; the readout compares the two norms, the two geodesic distances from $p$, and the chord distance between the two landing points.
A tangent step $v$ from $p$. Blue lands by the exponential map, magenta by add-and-renormalise, the hollow dot is the illegal $p+v$ floating off the sphere.
Why rotations do not add
SO(2) adds; SO(3) does not
In the plane, composing two rotations adds their angles and order does not matter: $R(\alpha)R(\beta) = R(\alpha+\beta)$, because $SO(2)$ is commutative. That is the exception. In three dimensions, rotating about $x$ and then about $y$ is not the same as the other order, and neither is a rotation by $\alpha+\beta$:
The leftover is the commutator of the two generators — the first term of the Baker–Campbell–Hausdorff series — so rotation parameters cannot be added like numbers. The correct picture is always: compose in the group (multiply matrices), and perturb in the tangent space (the Lie algebra $\mathfrak{so}(3)$). We do not re-derive Rodrigues' formula or the matrix exponential here; the rotation primer linked at the end builds both from scratch. The canvas is the friendly $SO(2)$ case, where the two hands coincide exactly; the readout also tracks the $SO(3)$ analogue computed in the standard axis–angle representation.
In $SO(2)$ the composed hand (magenta) sits exactly on the dashed $\alpha+\beta$ hand. Toggle the order: nothing moves, because planar rotations commute — the readout shows where 3D stops agreeing.
Retractions drift
First-order accuracy is not enough over many steps
One step hides the difference; a loop exposes it. Say we want to follow a tangent field around the sphere — the flow that rotates about a fixed axis. Each step we evaluate the field, take a step of size $h$, and either exponentiate or renormalise. Both integrate the same ODE to first order, with local error $O(h^2)$ per step, so over $K$ steps the two trajectories separate by $O(Kh^2)$:
Press Play to walk both around the loop. The readout accumulates the geodesic separation step by step and reports how badly each trajectory fails to close. Enlarge the step and the two retractions part company faster — this is why production solvers use a retraction with the right geometry, not merely any projection.
Blue: exponential-map steps. Magenta: add-and-renormalise steps. Both chase the same axis-rotation field for one revolution.
Optimising on the manifold
Riemannian gradient descent
Now put the pieces together. To minimise $f$ on $M$: project the ambient gradient into the tangent space, step downhill there, and retract back onto the surface. That is Riemannian gradient descent, and it is the flat algorithm with one projection added:
The two runs below minimise $f(p) = 0.7\,p_z + 0.35\,p_x p_y$ on the same sphere from the same start, one exponentiating and one renormalising. The exact minimum is the south pole $p^\star = (0,0,-1)$, where $\nabla f = -0.7\,p^\star$ is purely normal and the tangential gradient vanishes. Slide the learning rate: both get there, but the retraction path peels away from the geodesic one. This is the machinery behind fitting rotations — rotations are a manifold too, and its tangent-space perturbation is built in the rotation primer.
Blue: exponential-map descent. Magenta: renormalised descent. The amber ring marks the minimum at the south pole.
Where this shows up
From a sphere to a rotation
The sphere is the toy; rotations are the job. Optimising a camera pose, a robot joint, or a neural network on $SO(3)$ uses exactly the three moves above — a tangent space, an exponential map or retraction, and a Riemannian gradient. The matrix exponential, Rodrigues' formula, the tangent-space perturbation $R\cdot\exp(\delta)$ and the Jacobian that drives Gauss–Newton on a rotation are all worked out in SO(3) From Scratch, and the full theory, including poses, the adjoint and uncertainty, is the Lie Groups & Lie Algebras course. The hill-walking itself — steps, momentum, Newton and trust regions, now on a curved surface — is Nonlinear Optimization, and the flat theory of why those steps behave as they do was Curvature and convergence. On the arm, the joint vector is a chart for a torus, and every "reach" is a descent on that manifold.
Notation to carry forward
| Object | Reads as | Where it comes up |
|---|---|---|
M, TpM | A manifold and its tangent space at p — the local Rn of directions | Every curved parameter space |
∇Mf = ∇f − (∇f·p)p | Riemannian gradient: the ambient gradient minus its normal part | The only legal descent direction on M |
expp(v) | Geodesic from p with initial velocity v, length |v| | Exact surface moves; rotations via the matrix exponential |
Rp(v) | A retraction — any first-order stand-in for exp, e.g. (p+v)/|p+v| | Cheap updates when exp is expensive |
geodesic vs chord | Arc distance acos(p·q) versus straight-line |p−q| | Measuring error without leaving the surface |
R(α)R(β) vs R(α+β) | Compose versus add angles — equal in SO(2), unequal in SO(3) | Why rotation parameters must not be added |
pk+1 = exppk(−η∇Mf) | Riemannian gradient descent | Learning on rotations, poses and the sphere |
Further reading
- Boumal, An Introduction to Optimization on Smooth Manifolds — retractions, the Riemannian gradient and Riemannian gradient descent, all in one free book.
- Absil, Mahony & Sepulchre, Optimization Algorithms on Matrix Manifolds — the reference for retractions and their convergence orders.
- Lee, Introduction to Smooth Manifolds — charts, tangent spaces and the exponential map treated properly.