Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

What makes a matrix a rotation?

Starting point

🎯 Learning goal: a rotation is any linear map that preserves lengths and angles and doesn't flip the space inside-out. Turn that sentence into two matrix conditions you can check by hand.

A 3×3 matrix R rotates vectors — it never stretches, shears, or mirrors them. "Preserves lengths and angles" is exactly the statement (Ru)·(Rv) = u·v for every pair of vectors u, v. Expand the dot product and that condition collapses to one matrix equation:

RᵀR = I    ("orthogonal": R's columns are unit length and mutually perpendicular)

That alone allows one more possibility than rotation: a reflection (flip one axis) also satisfies RᵀR = I. Reflections turn a right-handed coordinate frame into a left-handed one, which shows up as det(R) = -1 instead of +1. Ruling that out gives the actual definition used everywhere in robotics:

SO(3) = { R ∈ ℝ³×³ : RᵀR = I, det(R) = +1 }    the "special orthogonal group" — every 3D rotation, and nothing else

Build a rotation below with an axis and an angle (this is anticipating section 3 — for now just treat it as "some matrix generator") and watch both checks hold for every axis and every angle, while a plain "add 1 to a diagonal entry" matrix (a not-quite-rotation) fails them.

test vector v Rv

Go deeper → Groups, matrix groups and manifolds: the group axioms, the matrix-group zoo, and why SO(3) rather than O(3).

2

The cross product in matrix form

The building block

🎯 Learning goal: "rotate around axis k" is really "keep sweeping perpendicular to k," and sweeping perpendicular to a fixed vector is exactly what the cross product does. Writing that cross product as matrix multiplication is the one trick everything below is built on.

For a fixed unit vector k = (k₁, k₂, k₃), the map v ↦ k × v is linear in v — so it's some matrix times v. Write out the cross product component by component and that matrix falls out directly:

[k]ₓ = ⎡ 0  -k₃  k₂ ⎤ ⎢ k₃   0  -k₁ ⎥    so that    [k]ₓv = k × v  for every v ⎣ -k₂  k₁   0 ⎦

[k]ₓ (read "k hat" or "the skew of k") is skew-symmetric: flipping it over its diagonal negates it, [k]ₓᵀ = -[k]ₓ. That's not a coincidence — it's forced by k × v always being perpendicular to v, i.e. vᵀ(k×v) = 0 for every v, which is exactly what Aᵀ = -A means for a matrix. Pick a k below and see its skew matrix, and watch [k]ₓv come out perpendicular to v every time, no matter what v is.

axis k v k × v = [k]ₓv
[k]ₓ

Go deeper → The tangent space and so(3): hat and vee, angular velocity in two frames, and the Lie bracket.

3

Deriving Rodrigues' rotation formula

From geometry to a formula

🎯 Learning goal: split any vector into a part along the rotation axis (untouched) and a part perpendicular to it (the part that actually sweeps around in a circle), write that circular sweep with sin and cos, and reassemble — that's the whole derivation.

To rotate v by angle θ about unit axis k, split v into a piece parallel to k and a piece perpendicular to it:

v∥ = (k·v) k    (stays fixed — it's on the axis) v⊥ = v - v∥    (this is the part that sweeps in a circle around k)

The perpendicular part v⊥ and the vector k×v⊥ = k×v are perpendicular to each other and to k, and both have the same length — exactly the two "clock hand" directions of a circle in the plane through v⊥. Standard 2D rotation-by-θ in that plane, written back in 3D vectors, gives:

Rv = v∥ + cosθ · v⊥ + sinθ · (k × v)

Substitute v∥ = v - v⊥, replace every cross product with [k]ₓ from section 2 (so k×v = [k]ₓv, and k×(k×v) = [k]ₓ²v = v⊥ - v, i.e. v⊥ = v + [k]ₓ²v), and everything collects into one matrix acting on v — Rodrigues' rotation formula:

R = I + sinθ [k]ₓ + (1-cosθ) [k]ₓ²

Drag axis and angle below — the canvas shows v sweeping to Rv exactly along the circle this derivation describes, and the readout confirms it stays at constant radius from the axis.

axis k v Rv
R = I + sinθ[k]ₓ + (1-cosθ)[k]ₓ²

Go deeper → Exp and Log on SO(3): the logarithm, its numerical traps near 0 and π, and the ball of radius π.

4

Why it's called a matrix exponential

A second route to the same formula

🎯 Learning goal: the ordinary exponential eˈx is defined by its Taylor series. Plug a skew-symmetric matrix into that same series and, after using [k]ₓ³ = -[k]ₓ, it collapses into exactly the sin/cos formula from section 3 — that's why part 5 calls it Exp(δ).

The matrix exponential is defined the same way the scalar one is, as an infinite sum:

exp(A) = I + A + A²/2! + A³/3! + A⁴/4! + …

Set A = θ[k]ₓ for unit axis k. A short computation shows [k]ₓ³ = -[k]ₓ (a skew-symmetric matrix built from a unit vector cubes back to minus itself) — so every power beyond the square is just ±[k]ₓ or ±[k]ₓ², and the whole series sorts into two families:

exp(θ[k]ₓ) = I + (θ - θ³/3! + θ⁵/5! - …)[k]ₓ + (θ²/2! - θ⁴/4! + …)[k]ₓ² = I + sinθ [k]ₓ + (1-cosθ) [k]ₓ²    — the two parenthesized series are exactly sinθ and (1-cosθ)

Same formula as section 3, reached a completely different way. This is the function part 5 calls Exp(δ): feed it a 3-number vector δ (its direction is the axis k, its length is the angle θ), and it hands back a rotation matrix. Pick a truncation order below and watch the truncated series converge onto the exact matrix as you add more terms — by 3-4 terms it's already indistinguishable for a moderate angle.

⚠️ Code's expMap(w) in part 5 doesn't sum a series at runtime — it plugs θ = |w| and k = w/|w| straight into the closed-form I + sinθ[k]ₓ + (1-cosθ)[k]ₓ². The series is only here to show why that closed form deserves the name "exponential."

Go deeper → Exp and Log on SO(3): the series derivation with a convergence chart for any angle.

5

Why perturbations compose instead of add

The step that made part 5's update rule work

🎯 Learning goal: SO(3) is a curved set of matrices, not a flat space — you can't leave it by adding two rotations together, only by composing them. Exp(δ) is the tool that turns a flat 3-number "direction to move" into a valid step that stays on SO(3).

Adding two arbitrary rotation matrices entry-by-entry almost never gives another rotation — the sum generally fails RᵀR=I (try it: I + I = 2I, and (2I)ᵀ(2I) = 4I ≠ I). So "current estimate plus a small step" can't mean ordinary addition the way it does for a position in ℝ³. It has to mean compose — multiply on a small extra rotation:

Rₕ₊₁ = Rₕ · Exp(δ)    δ ∈ ℝ³ is a flat 3-number "tangent" direction; Exp bends it onto SO(3)

Because Exp is built from an exact sin/cos formula, R·Exp(δ) is a genuine rotation matrix for any size of δ, not just small ones — there's no approximation being made by using it. The approximation people usually reach for is the opposite direction: for a small δ, the exponential series' higher terms are negligible, so Exp(δ) ≈ I + [δ]ₓ — which is exactly why an optimizer's tiny per-iteration step can be treated as approximately linear even though the underlying space is curved. That local flatness is what makes gradient descent, Gauss-Newton and Levenberg-Marquardt (all of which reason about a linear step) work at all on something curved like SO(3).

Drag δ below and compare the exact composed rotation to the linearized I+[δ]ₓ stand-in (renormalized back onto SO(3) so it's a fair comparison) — small δ, they agree; push it larger and the linear stand-in visibly drifts off, exactly the same shape of error part 5 showed for "naive" angle-adding.

exact: R·Exp(δ) linearized: R·normalize(I+[δ]ₓ)

Go deeper → SO(2): the whole story in one dimension: ⊕ and ⊖ where every formula fits on a line.

6

Estimating the Jacobian in the tangent space

How the code actually computes a step

🎯 Learning goal: a Jacobian column is "how much does the residual change per unit nudge along one tangent direction." With no formula for that derivative on hand, central differences get a numerical answer by nudging forward and backward in the tangent space and dividing by the step.

For a residual r(R) = measured - Rᵀn comparing a landmark direction rotated into the sensor frame against what was actually measured, part 5's Gauss-Newton needs ∂r/∂δ — but δ isn't a coordinate R is written in terms of, it's the tangent nudge from section 5. The fix: physically apply a tiny nudge in each of the 3 tangent directions, using the exact same R·Exp(δ) composition, and see how much r moved:

column k of J  ≈  [ r(R·Exp(h·eₖ)) − r(R·Exp(-h·eₖ)) ] / (2h)    eₖ = k-th standard basis vector, h ≈ 1e-6

That's a literal transcription of gradAndHessNumeric in part 5's script: three tiny nudges (one per tangent axis), each composed on with Exp, each evaluated forward and backward, divided by 2h. Pick a landmark and a tangent axis below and watch this play out on the exact numbers — as the step size h shrinks from coarse to tiny, the estimate stabilizes onto essentially one value (until it's so tiny that floating-point rounding starts to dominate, visible on the far left of the slider).

Go deeper → Jacobians on Lie groups: the analytic right Jacobian that replaces these central differences, with a checkable table of elementary derivatives.

7

Assembling the step: back to Gauss-Newton

Where this rejoins part 5

🎯 Learning goal: once J and the gradient g = Jᵀr are in hand, the normal equations from parts 1–4 don't change at all — only what the resulting δ gets applied to changes.

With the numerical J from section 6 (stacked over all landmarks) and gradient g = Jᵀr, Gauss-Newton solves the exact same 3×3 linear system as every earlier part — (JᵀJ) δ = -g — for the tangent step δ. Levenberg-Marquardt damps it the same way too, adding λI before solving. The only thing that's new versus parts 1–4 is the last line:

solve (JᵀJ) δ = -g  for δ ∈ ℝ³    // identical to a 2D or nD least-squares step R ← R · Exp(δ)                // the one line that's different: compose, don't add

That's the entire gap this page fills in: sections 1-2 built the vocabulary (what a rotation is, and the skew-symmetric matrix that turns a cross product into multiplication), sections 3-4 derived Exp two ways, section 5 explained why composing is mandatory and when the linear approximation is safe, and section 6 showed the Jacobian actually being computed with it. Everything downstream — the normal equations, the damping, the convergence loop — is unchanged from part 1.

Go deeper → Optimization on manifolds: Gauss–Newton on SO(3) and SE(3), rotation averaging and a pose graph.

✓

Check yourself

Five quick questions

✓

Cheat sheet

Recap

ObjectWhat it is
SO(3)All 3×3 matrices R with RᵀR=I and det(R)=+1 — every possible 3D rotation, nothing else
[k]ₓSkew-symmetric matrix with [k]ₓv = k×v; the matrix form of "rotate around k, infinitesimally"
Rodrigues' formulaR = I + sinθ[k]ₓ + (1-cosθ)[k]ₓ² — an axis k and angle θ turned into an explicit rotation matrix
Exp(δ)Same formula, reparametrized by δ = θk (axis direction and angle folded into one 3-vector) — the matrix exponential of [δ]ₓ
Tangent stepR ← R·Exp(δ), never R + δ — composing keeps the result a valid rotation for any size of δ
Numerical JacobianCentral differences taken by nudging with Exp(±h·eₖ) along each tangent axis, not by differentiating R's nine entries directly
The step itselfUnchanged from part 1: (JᵀJ)δ = -Jᵀr, optionally damped by λI — only what δ updates is new
Ready to see this machinery running the full demos? → Back to part 5