Rotations, SO(3) and the exponential map
Every rotation of three-dimensional space preserves lengths and angles, and there is exactly one class of matrix that does that without reflecting: an orthonormal matrix with determinant +1. Those matrices form a group called SO(3), and it is not a vector space — you cannot add two rotations and expect a rotation, and a straight line between two of them leaves the set entirely. This part shows the two coordinates on SO(3) that actually work: an angle and a unit axis, which is the matrix exponential of a skew matrix, and quaternions, which are the same object in four numbers. It also shows where the familiar Euler angles fail, and why the fix is to move along a curve that stays on the manifold.
The question
Why a rotation is a curve, not a vector
By now you have met matrices as transformations and eigenvectors as the directions they leave alone. A rotation is a transformation with no preferred direction at all: there is no real eigenvector with eigenvalue 1 pointing somewhere interesting, because every vector moves. What a rotation does keep is much stronger than any single direction. It keeps every length and every angle, which is what it means to move rigidly. That single requirement pins the matrix down almost completely.
The catch is that the set of all rotations is curved. Take two rotations, average their matrices entry by entry, and the average is almost never a rotation — it shears, it scales, it can even collapse the cube you were turning. So everything that works for vectors stops working here: no straight-line interpolation, no simple subtraction of one orientation from another, no ordinary calculus of the form x(t) = x₀ + t v. The story of this part is how to put coordinates on a curved set so that the ordinary machinery can be recovered.
Rotations are orthonormal matrices with determinant +1
The definition that does all the work
Write the rotation as a matrix R whose columns are the images of the three basis vectors. Preserving length and angle is the statement that the columns stay unit length and stay mutually perpendicular, which is exactly the pair of conditions RᵀR = I. A matrix with that property is called orthogonal, and its columns form an orthonormal frame — an orientation of space. The three rotated axes are still a rigid tripod, just pointed somewhere new.
That determinant is the last bit of freedom. Orthogonality forces |det R| = 1, so the determinant is either +1 or −1. The −1 case is a rotation composed with a mirror: it preserves lengths, but it turns a right hand into a left hand. The proper rotations, the ones a rigid body can actually perform, are the ones with det R = +1. This set is closed under multiplication, contains the identity, and every element has an inverse in the set; mathematicians call it the special orthogonal group, SO(3). The word group matters because it says the composition of rotations is again a rotation, while nothing guarantees that a sum is.
You can check both identities numerically on any rotation you like in the demo below: the readout keeps RᵀR within a hair of the identity and the determinant within a hair of +1, no matter how you drag the axis and the angle.
One warning before you go further: rotations do not commute. Turn a book a quarter turn about the vertical, then a quarter turn about the axis pointing away from you, and note where the spine ends up. Now do the two turns in the opposite order. The book finishes in a different orientation both times, and no amount of care with the numbers changes that. So the composition of rotations is a non-abelian group operation, and the habit from earlier parts of treating a transformation as a matrix you can reorder has to be set aside. This non-commutativity is the reason a sequence of Euler angles depends on its order, and it is the reason the logarithm of a product of rotations is not the sum of their logarithms.
Spin by θ about a unit axis û
Euler's theorem, Rodrigues' formula, and the exponential map
Euler's rotation theorem says something surprising for a curved, three-parameter family: every single rotation, no matter how it was built from a sequence of turns, is a turn about one fixed axis. Pick a unit vector û — the direction that does not move — and an angle θ, and you have named the rotation completely. The axis is an eigenvector with eigenvalue 1, and the angle is how far the plane perpendicular to it turns. This angle-axis description is the first coordinate system on SO(3) that never lies to you about what is happening.
To turn that sentence into a matrix, write the cross-product operation as a matrix. The skew-symmetric matrix [û]× is defined by [û]× v = û × v, and it is the infinitesimal generator of rotations about û. Compose it with itself and the pattern closes: successive powers only ever return to K or K² with alternating signs, so the familiar infinite series for the matrix exponential collapses to three terms. This is Rodrigues' formula, and it is the exponential map from the Lie algebra so(3) to the group SO(3):
Read the formula as a recipe. The identity does nothing; the K term is the first-order spin; the K² term is the correction that keeps the result orthonormal for large angles. Because the axis is a unit vector, only two numbers are needed to choose it and one to choose the angle — three numbers for a three-dimensional group, which is why this is a good coordinate system and not just a convenient one. The demo below builds R exactly this way, with the axis set by two sliders and the angle by a third.
The magenta arrow is the axis û; the cube is the standard basis cube turned by R = exp(θ[û]×). Drag the axis angles and the spin angle.
A quaternion is the same rotation in four numbers. From the axis and angle, form q = (cos(θ/2), sin(θ/2) û) — one scalar part and a three-vector part. The half angle is not a trick: applying q twice composes the rotation twice, which is exactly why half angles appear. The label below shows the quaternion alongside the matrix, and it reveals the double cover: q and −q name the same rotation, so the quaternion sphere wraps the space of rotations twice. That redundancy is also the reason quaternions never hit a coordinate singularity, and it is why robotics code usually stores orientation as a unit quaternion and only converts to a matrix once per frame.
Euler angles and gimbal lock
Three familiar sliders that collide
The other way to name an orientation is the one every flight sim and robot arm uses: three angles applied in a fixed order, say yaw about the world z, then pitch about the new y, then roll about the newest x, giving R = R_z(γ) R_y(β) R_x(α). Each slider feels independent, which is why the representation is so popular. The problem is that the three axes are not fixed in space; each rotation carries the next axis with it, and for one particular pitch the carried axis lands exactly on top of the first one.
Rotate the pitch to ±90° in the demo and watch. The roll axis, which started as the body x direction, ends up parallel to the world z axis that yaw already rotates about. Two of the three sliders now turn the body about the same line, so their effects add and subtract rather than acting independently: a whole degree of freedom has been swallowed, and the mapping from angles to rotations loses rank. This is gimbal lock. It is not a numerical accident; it is a genuine singularity of the coordinates, the place where the chart on SO(3) folds over.
Faint arrows are the fixed world axes; bright arrows are the body axes. The magenta body x axis is the roll axis. Push pitch to ±90° and it lines up with the highlighted world z.
Notice that no cube ever becomes non-rigid here — every value of the three sliders still produces a perfectly good rotation, because the product of rotations is always a rotation. What breaks is the ability to move smoothly through orientations using the three angles as coordinates. Near the lock, a small change in roll and a small change in yaw produce almost the same motion, so a controller that differentiates with respect to the angles sees a nearly singular Jacobian and commands enormous, jittery corrections. The cure is to leave Euler angles for anything that has to be interpolated, differentiated or optimised, and use the axis-angle or quaternion coordinates instead.
It helps to see the singularity as a rank statement rather than a drawing. The derivative of the orientation with respect to the three angles is a 3×3 matrix built from the three instantaneous rotation axes; when two of those axes become parallel, two of its columns become parallel too, and the matrix drops rank. At that point the map from angle-space to SO(3) is locally two-to-one in one direction, which is why an inverse-kinematics or attitude solver stalls. Every three-angle convention has such points — changing the order or the axis set only moves them around, it never removes them, because three numbers cannot smoothly parametrise a curved three-dimensional set without at least one fold.
Four ways to name the same orientation
| Coordinates | Size | Strength | Where it fails |
|---|---|---|---|
| Rotation matrix | 9 numbers, 6 constraints | Compose and apply directly; no trigonometry in the hot loop | Over-parametrised; drifts off SO(3) under floating-point error |
| Euler angles | 3 numbers | Human-readable; matches how gimbals and joints are physically built | Gimbal lock at the singular pitch; interpolation is not a geodesic |
| Axis-angle | 3 numbers plus a unit axis | Exactly the exponential map; every rotation has one | The axis is undefined at θ = 0; angle wraps at 2π |
| Unit quaternion | 4 numbers, 1 constraint | No singularity; slerp is the geodesic and is cheap | Double cover: q and −q are the same rotation, so signs must be normalised |
As a rule of thumb: store the state as a matrix or a quaternion, expose Euler angles only at the user interface, and do every increment, derivative and interpolation in the tangent space with an exponential and a logarithm.
exp, log and interpolation on the manifold
Staying on SO(3) instead of cutting through it
To move between two orientations, you first need the logarithm, the inverse of the exponential map. Given a rotation R, its trace gives the angle through tr R = 1 + 2 cos θ, and the antisymmetric part gives the axis: the vector (R − Rᵀ) is 2 sin θ times the skew matrix of the axis. Together they recover the pair (θ, û) that the exponential map consumes. Composing the two maps, log(R₀ᵀR₁) gives the single rotation that carries one orientation to the other, as an axis and an angle.
Now the geodesic is easy to write down. Take a fraction t of that relative angle along the same axis and rebuild a rotation: R(t) = R₀ \exp(t \log(R₀ᵀR₁)). Every value of t is a genuine rotation, the cube turns rigidly, and the path is the shortest one on the manifold — the analogue of a straight line, bent to live on the curved set. Averaging the matrices directly, (1-t)R₀ + tR₁, does something else entirely: it takes a chord through the interior, and the result has determinant less than one and columns that are no longer orthogonal, so the cube shears and shrinks. The demo draws both, so you can see the rotation and the impostor at the same time.
Solid cube: the geodesic R(t). Faded cube: the entrywise average, which is not a rotation. Slide t from one orientation to the other.
This is exactly what spherical linear interpolation, slerp, computes when it is written with quaternions instead of matrices: shortest-arc interpolation between two unit quaternions traces the same geodesic. The quaternion version is cheap, has no 90° singularity, and chooses the short way around by flipping the sign of one quaternion when their dot product is negative. Matrix log and quaternion slerp are two implementations of one geometric idea, and the exponential map is the bridge between them.
There is a second reason the exponential map is the right primitive, and it is the one optimisation code cares about. Around any orientation, the tangent space at that point is a genuine vector space: small increments can be added, scaled, and solved for with the linear-algebra machinery of the rest of this series. A solver keeps its estimate on the manifold and its error in the tangent space, updates the estimate by R ← R \exp(\delta), and recomputes the local linearisation at the new point. That single pattern — retract with an exponential, linearise in the tangent space, repeat — is what makes pose-graph optimisation, bundle adjustment and attitude estimation tractable despite the curvature. The non-commutativity from earlier shows up here as a correction term in the composition of two exponentials, which is small when the increments are small, and is exactly why these methods iterate rather than solve in one shot.
Where this shows up
One manifold, two worlds
Attitude estimation and pose on the manifold
A robot's orientation is an element of SO(3), and every good estimator keeps it as one — a unit quaternion or a rotation matrix updated by an exponential of an increment — rather than as Euler angles that can lock. The increments themselves are tangent vectors, and optimising over them is what makes pose graphs and SLAM solvable. See 3D rotations for the geometry and SO(3) from scratch for the full derivation.
Symmetry and flows on curved spaces
A model that respects rotation is built to commute with SO(3), which cuts the parameters it needs to learn; the same exponentials appear in normalising flows whose densities live on manifolds rather than in flat space. The architecture chapter shows how geometry and symmetry shape the layers of a modern network, and the group language here is what those constructions are written in.
Further reading
- Grant Sanderson, "Abstract vector spaces", Essence of Linear Algebra, 3Blue1Brown — the chapter that makes the jump from arrows to sets closed under the same operations, the frame this part uses for a group.
- Gilbert Strang, 18.06 Linear Algebra, MIT OpenCourseWare — the lectures on orthogonal matrices and the singular-value decomposition; Lecture 21 and the following lectures are where the orthonormal frame is developed.
- Denim Patel, Optimization with rotations and SO(3) from scratch — the applied companion pages where these matrices meet real estimators.
- Lie Groups & Lie Algebras, Interactively — a 12-part course that takes this page's exponential map all the way to SE(3), Jacobians, uncertainty and IMU preintegration.
- Joan Solà, Quaternion kinematics for the error-state Kalman filter (arXiv) — the reference for treating orientation as a manifold with a tangent-space error, the pattern behind most modern attitude filters.