Rotations Beyond a Single Angle
← Part 4 recap: trajectories, drift, and loop closure
Every part so far lived in a flat 2D plane, where a heading was just one number, θ, that wrapped around every 360°. Real robots, drones, and cameras rotate in full 3D — and in 3D, orientation stops being one number, stops commuting, and stops being something you can just add to. This last part is about that one shift in how rotation itself behaves, and shows that gradient descent, Newton, Gauss-Newton and Levenberg-Marquardt still solve it once you know the right way to take a step.
What's new: orientation isn't one number anymore
Recap
In 2D, "which way am I facing" is a single angle θ, and the only wrinkle was that it wraps around (parts 2–4). In 3D, orientation needs three independent numbers — think roll, pitch, and yaw — but those three numbers behave nothing like an (x, y, z) position. You can't average two orientations by averaging their numbers, and — this is the genuinely new part — the order you apply two rotations in changes the result.
Rotations don't commute
The surprise
Both triads below start pointing the same way and get the exact same two rotations applied — one around the x-axis, one around the y-axis — just in the opposite order. Drag the sliders and watch them end up pointing in genuinely different directions. This is why 3D orientation can't be treated as a vector of three independent numbers you add together.
Perturbing a rotation: compose, don't add
The fix
Exp(δ) — the matrix exponential of a small rotation vector, also called Rodrigues' formula — turns that 3-number nudge into an actual tiny rotation matrix, which then gets composed onto the current orientation. Drag the nudge sliders below: the solid triad applies the nudge correctly, by composing; the faint dashed one applies the same numbers naively, just adding them to the equivalent roll/pitch/yaw angles. For a tiny nudge they're nearly identical — push the sliders further and watch them visibly diverge.
Exp(δ), skew-symmetric matrices, or why composing is mandatory here? The math primer derives Rodrigues' formula from scratch, two different ways, with its own interactive widgets — worth a detour before continuing. For the complete theory (quaternions, SE(3), the adjoint, analytic Jacobians and uncertainty), see the 12-part Lie Groups & Lie Algebras course.Finding an orientation from landmark sightings
Gauss-Newton, in 3D
Four landmarks, each observed as a noisy 3D direction in the robot's own frame. Step through Gauss-Newton (its Jacobian here is estimated numerically rather than hand-derived — deriving 3D rotation derivatives analytically needs more differential-geometry machinery than fits on this page, but the update rule is identical either way) and watch the colored dots — each one a landmark direction, rotated into the world by the current orientation estimate — slide onto the fixed true directions.
Levenberg-Marquardt on the rotation
Same damping, one more time
λI to a 3×3 JᵀJ — the dimensionality of the tangent nudge, not the representation of the rotation itself, is all that ever mattered.Same landmarks, but this time starting almost 180° off — pointing roughly the opposite way. Gauss-Newton's numerical Jacobian is a much rougher local model out here; adaptive λ keeps the early steps honest.
Playground: race all three, in 3D
Put it all together
Same four-landmark scenario, same badly-wrong starting orientation, three methods racing at once (Newton is left out here too, for the same reason as parts 3–4). The table tracks the remaining angle between the estimate and the true orientation.
| Method | Iters | Angle err | Status |
|---|
Cheat sheet — and the whole series
Recap
| Concept | What changed from part 4 |
|---|---|
| Unknown | A 2D heading θ → a full 3D orientation (3 degrees of freedom, not 1) |
| New surprise | Rotations don't commute — order matters, unlike adding angles in 2D |
| Storage vs. update | Store orientation as a matrix or quaternion; update it with a small 3-number tangent nudge, composed via R · Exp(δ) |
| Jacobian | Estimated numerically here, rather than hand-derived — same idea, less differential geometry |
| Update rules | Identical in spirit to every part before this — gradient descent, Gauss-Newton, and Levenberg-Marquardt all still apply, over whatever the tangent space happens to be |
Across five parts, one method kept reappearing in a new costume every time: turn "what's the right answer" into "what minimizes a cost function," take the best step you can with the information available — slope alone (gradient descent), slope and curvature (Newton), a cheap stand-in for curvature (Gauss-Newton), or a damped blend of the two (Levenberg-Marquardt) — and repeat. Position, then heading, then an unknown map, then a whole trajectory, then a rotation that doesn't even live in ordinary space: the update rule barely changed. That's the actual lesson underneath all the robot-localization dressing — nonlinear least-squares is one idea, reused everywhere.