SO(3) From Scratch
← Back to part 5: Rotations Beyond a Single Angle
Part 5 used three facts about 3D rotations without deriving them: that a small nudge composes as R·Exp(δ) instead of adding, that Exp(δ) is "the matrix exponential, also called Rodrigues' formula," and that the Jacobian can be estimated numerically inside that same formula. This page is the derivation those facts skipped — built up from ordinary vectors and matrices, with nothing assumed beyond first-year linear algebra. Every widget below computes with the exact same functions the part 5 demos use.
What makes a matrix a rotation?
Starting point
A 3×3 matrix R rotates vectors — it never stretches, shears, or mirrors them. "Preserves lengths and angles" is exactly the statement (Ru)·(Rv) = u·v for every pair of vectors u, v. Expand the dot product and that condition collapses to one matrix equation:
That alone allows one more possibility than rotation: a reflection (flip one axis) also satisfies RᵀR = I. Reflections turn a right-handed coordinate frame into a left-handed one, which shows up as det(R) = -1 instead of +1. Ruling that out gives the actual definition used everywhere in robotics:
Build a rotation below with an axis and an angle (this is anticipating section 3 — for now just treat it as "some matrix generator") and watch both checks hold for every axis and every angle, while a plain "add 1 to a diagonal entry" matrix (a not-quite-rotation) fails them.
Go deeper → Groups, matrix groups and manifolds: the group axioms, the matrix-group zoo, and why SO(3) rather than O(3).
The cross product in matrix form
The building block
For a fixed unit vector k = (k₁, k₂, k₃), the map v ↦ k × v is linear in v — so it's some matrix times v. Write out the cross product component by component and that matrix falls out directly:
[k]ₓ (read "k hat" or "the skew of k") is skew-symmetric: flipping it over its diagonal negates it, [k]ₓᵀ = -[k]ₓ. That's not a coincidence — it's forced by k × v always being perpendicular to v, i.e. vᵀ(k×v) = 0 for every v, which is exactly what Aᵀ = -A means for a matrix. Pick a k below and see its skew matrix, and watch [k]ₓv come out perpendicular to v every time, no matter what v is.
Go deeper → The tangent space and so(3): hat and vee, angular velocity in two frames, and the Lie bracket.
Deriving Rodrigues' rotation formula
From geometry to a formula
To rotate v by angle θ about unit axis k, split v into a piece parallel to k and a piece perpendicular to it:
The perpendicular part v⊥ and the vector k×v⊥ = k×v are perpendicular to each other and to k, and both have the same length — exactly the two "clock hand" directions of a circle in the plane through v⊥. Standard 2D rotation-by-θ in that plane, written back in 3D vectors, gives:
Substitute v∥ = v - v⊥, replace every cross product with [k]ₓ from section 2 (so k×v = [k]ₓv, and k×(k×v) = [k]ₓ²v = v⊥ - v, i.e. v⊥ = v + [k]ₓ²v), and everything collects into one matrix acting on v — Rodrigues' rotation formula:
Drag axis and angle below — the canvas shows v sweeping to Rv exactly along the circle this derivation describes, and the readout confirms it stays at constant radius from the axis.
Go deeper → Exp and Log on SO(3): the logarithm, its numerical traps near 0 and π, and the ball of radius π.
Why it's called a matrix exponential
A second route to the same formula
Exp(δ).The matrix exponential is defined the same way the scalar one is, as an infinite sum:
Set A = θ[k]ₓ for unit axis k. A short computation shows [k]ₓ³ = -[k]ₓ (a skew-symmetric matrix built from a unit vector cubes back to minus itself) — so every power beyond the square is just ±[k]ₓ or ±[k]ₓ², and the whole series sorts into two families:
Same formula as section 3, reached a completely different way. This is the function part 5 calls Exp(δ): feed it a 3-number vector δ (its direction is the axis k, its length is the angle θ), and it hands back a rotation matrix. Pick a truncation order below and watch the truncated series converge onto the exact matrix as you add more terms — by 3-4 terms it's already indistinguishable for a moderate angle.
expMap(w) in part 5 doesn't sum a series at runtime — it plugs θ = |w| and k = w/|w| straight into the closed-form I + sinθ[k]ₓ + (1-cosθ)[k]ₓ². The series is only here to show why that closed form deserves the name "exponential."Go deeper → Exp and Log on SO(3): the series derivation with a convergence chart for any angle.
Why perturbations compose instead of add
The step that made part 5's update rule work
SO(3) is a curved set of matrices, not a flat space — you can't leave it by adding two rotations together, only by composing them. Exp(δ) is the tool that turns a flat 3-number "direction to move" into a valid step that stays on SO(3).Adding two arbitrary rotation matrices entry-by-entry almost never gives another rotation — the sum generally fails RᵀR=I (try it: I + I = 2I, and (2I)ᵀ(2I) = 4I ≠ I). So "current estimate plus a small step" can't mean ordinary addition the way it does for a position in ℝ³. It has to mean compose — multiply on a small extra rotation:
Because Exp is built from an exact sin/cos formula, R·Exp(δ) is a genuine rotation matrix for any size of δ, not just small ones — there's no approximation being made by using it. The approximation people usually reach for is the opposite direction: for a small δ, the exponential series' higher terms are negligible, so Exp(δ) ≈ I + [δ]ₓ — which is exactly why an optimizer's tiny per-iteration step can be treated as approximately linear even though the underlying space is curved. That local flatness is what makes gradient descent, Gauss-Newton and Levenberg-Marquardt (all of which reason about a linear step) work at all on something curved like SO(3).
Drag δ below and compare the exact composed rotation to the linearized I+[δ]ₓ stand-in (renormalized back onto SO(3) so it's a fair comparison) — small δ, they agree; push it larger and the linear stand-in visibly drifts off, exactly the same shape of error part 5 showed for "naive" angle-adding.
Go deeper → SO(2): the whole story in one dimension: ⊕ and ⊖ where every formula fits on a line.
Estimating the Jacobian in the tangent space
How the code actually computes a step
For a residual r(R) = measured - Rᵀn comparing a landmark direction rotated into the sensor frame against what was actually measured, part 5's Gauss-Newton needs ∂r/∂δ — but δ isn't a coordinate R is written in terms of, it's the tangent nudge from section 5. The fix: physically apply a tiny nudge in each of the 3 tangent directions, using the exact same R·Exp(δ) composition, and see how much r moved:
That's a literal transcription of gradAndHessNumeric in part 5's script: three tiny nudges (one per tangent axis), each composed on with Exp, each evaluated forward and backward, divided by 2h. Pick a landmark and a tangent axis below and watch this play out on the exact numbers — as the step size h shrinks from coarse to tiny, the estimate stabilizes onto essentially one value (until it's so tiny that floating-point rounding starts to dominate, visible on the far left of the slider).
Go deeper → Jacobians on Lie groups: the analytic right Jacobian that replaces these central differences, with a checkable table of elementary derivatives.
Assembling the step: back to Gauss-Newton
Where this rejoins part 5
With the numerical J from section 6 (stacked over all landmarks) and gradient g = Jᵀr, Gauss-Newton solves the exact same 3×3 linear system as every earlier part — (JᵀJ) δ = -g — for the tangent step δ. Levenberg-Marquardt damps it the same way too, adding λI before solving. The only thing that's new versus parts 1–4 is the last line:
That's the entire gap this page fills in: sections 1-2 built the vocabulary (what a rotation is, and the skew-symmetric matrix that turns a cross product into multiplication), sections 3-4 derived Exp two ways, section 5 explained why composing is mandatory and when the linear approximation is safe, and section 6 showed the Jacobian actually being computed with it. Everything downstream — the normal equations, the damping, the convergence loop — is unchanged from part 1.
Go deeper → Optimization on manifolds: Gauss–Newton on SO(3) and SE(3), rotation averaging and a pose graph.
Check yourself
Five quick questions
Cheat sheet
Recap
| Object | What it is |
|---|---|
SO(3) | All 3×3 matrices R with RᵀR=I and det(R)=+1 — every possible 3D rotation, nothing else |
[k]ₓ | Skew-symmetric matrix with [k]ₓv = k×v; the matrix form of "rotate around k, infinitesimally" |
| Rodrigues' formula | R = I + sinθ[k]ₓ + (1-cosθ)[k]ₓ² — an axis k and angle θ turned into an explicit rotation matrix |
Exp(δ) | Same formula, reparametrized by δ = θk (axis direction and angle folded into one 3-vector) — the matrix exponential of [δ]ₓ |
| Tangent step | R ← R·Exp(δ), never R + δ — composing keeps the result a valid rotation for any size of δ |
| Numerical Jacobian | Central differences taken by nudging with Exp(±h·eₖ) along each tangent axis, not by differentiating R's nine entries directly |
| The step itself | Unchanged from part 1: (JᵀJ)δ = -Jᵀr, optionally damped by λI — only what δ updates is new |