Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

0

What's new: orientation isn't one number anymore

Recap

In 2D, "which way am I facing" is a single angle θ, and the only wrinkle was that it wraps around (parts 2–4). In 3D, orientation needs three independent numbers — think roll, pitch, and yaw — but those three numbers behave nothing like an (x, y, z) position. You can't average two orientations by averaging their numbers, and — this is the genuinely new part — the order you apply two rotations in changes the result.

💡 The plan: see rotations fail to commute, learn the one rule for perturbing an orientation correctly, then point the exact same Gauss-Newton and Levenberg-Marquardt machinery from parts 1–4 at a 3D problem.
1

Rotations don't commute

The surprise

🎯 Learning goal: in 2D, stacking two rotations just adds two angles — order never mattered. In 3D, rotating around the x-axis then the y-axis lands somewhere completely different than y-axis-then-x-axis, for the exact same two rotations.

Both triads below start pointing the same way and get the exact same two rotations applied — one around the x-axis, one around the y-axis — just in the opposite order. Drag the sliders and watch them end up pointing in genuinely different directions. This is why 3D orientation can't be treated as a vector of three independent numbers you add together.

Rotate x, then y Rotate y, then x
2

Perturbing a rotation: compose, don't add

The fix

🎯 Learning goal: just like part 2's wrapped angle wasn't a plain subtraction, a small nudge to a 3D orientation isn't a plain addition either. The correct move is to compose a small rotation on top of the current one.
Rₕ₊₁ = Rₕ · Exp(δ)    δ = a small 3-number "nudge" (an axis to tilt around, and by how much)

Exp(δ) — the matrix exponential of a small rotation vector, also called Rodrigues' formula — turns that 3-number nudge into an actual tiny rotation matrix, which then gets composed onto the current orientation. Drag the nudge sliders below: the solid triad applies the nudge correctly, by composing; the faint dashed one applies the same numbers naively, just adding them to the equivalent roll/pitch/yaw angles. For a tiny nudge they're nearly identical — push the sliders further and watch them visibly diverge.

📖 New to Exp(δ), skew-symmetric matrices, or why composing is mandatory here? The math primer derives Rodrigues' formula from scratch, two different ways, with its own interactive widgets — worth a detour before continuing. For the complete theory (quaternions, SE(3), the adjoint, analytic Jacobians and uncertainty), see the 12-part Lie Groups & Lie Algebras course.
Composed: R · Exp(δ) Naive: angles added directly
⚠️ In practice: this is also why orientation is usually stored as a rotation matrix or a 4-number quaternion (which avoids the singularities plain roll/pitch/yaw angles run into — "gimbal lock") but updated during optimization using a 3-number tangent nudge like δ above. Storage representation and update representation don't have to be the same thing.
3

Finding an orientation from landmark sightings

Gauss-Newton, in 3D

🎯 Learning goal: the robot's position is known this time — only its orientation is unknown. Each landmark gives a noisy bearing (a 3D direction, not just an angle), and the residual is simply how far off that direction is from what the current orientation estimate predicts.
r⁢(R) = measured⁢ − Rᵀn⁢    n⁢ = true world-frame direction to landmark i (known)

Four landmarks, each observed as a noisy 3D direction in the robot's own frame. Step through Gauss-Newton (its Jacobian here is estimated numerically rather than hand-derived — deriving 3D rotation derivatives analytically needs more differential-geometry machinery than fits on this page, but the update rule is identical either way) and watch the colored dots — each one a landmark direction, rotated into the world by the current orientation estimate — slide onto the fixed true directions.

◆ true landmark direction estimate's predicted direction
4

Levenberg-Marquardt on the rotation

Same damping, one more time

🎯 Learning goal: the damping term still adds λI to a 3×3 JᵀJ — the dimensionality of the tangent nudge, not the representation of the rotation itself, is all that ever mattered.

Same landmarks, but this time starting almost 180° off — pointing roughly the opposite way. Gauss-Newton's numerical Jacobian is a much rougher local model out here; adaptive λ keeps the early steps honest.

5

Playground: race all three, in 3D

Put it all together

Same four-landmark scenario, same badly-wrong starting orientation, three methods racing at once (Newton is left out here too, for the same reason as parts 3–4). The table tracks the remaining angle between the estimate and the true orientation.

Gradient descent Gauss-Newton Levenberg-Marquardt
MethodItersAngle errStatus
✓

Cheat sheet — and the whole series

Recap

ConceptWhat changed from part 4
UnknownA 2D heading θ → a full 3D orientation (3 degrees of freedom, not 1)
New surpriseRotations don't commute — order matters, unlike adding angles in 2D
Storage vs. updateStore orientation as a matrix or quaternion; update it with a small 3-number tangent nudge, composed via R · Exp(δ)
JacobianEstimated numerically here, rather than hand-derived — same idea, less differential geometry
Update rulesIdentical in spirit to every part before this — gradient descent, Gauss-Newton, and Levenberg-Marquardt all still apply, over whatever the tangent space happens to be

Across five parts, one method kept reappearing in a new costume every time: turn "what's the right answer" into "what minimizes a cost function," take the best step you can with the information available — slope alone (gradient descent), slope and curvature (Newton), a cheap stand-in for curvature (Gauss-Newton), or a damped blend of the two (Levenberg-Marquardt) — and repeat. Position, then heading, then an unknown map, then a whole trajectory, then a rotation that doesn't even live in ordinary space: the update rule barely changed. That's the actual lesson underneath all the robot-localization dressing — nonlinear least-squares is one idea, reused everywhere.

Missed the earlier parts? Start with position-only localization. ← Back to part 1