Multi-View Geometry, Interactively
Multi-view geometry is the math behind turning flat photographs into 3D structure — how self-driving cars, AR headsets and 3D scanners figure out where things are from pictures alone. This guide builds it up one idea at a time, using one running scene: a small wireframe cube, watched by one or more pinhole cameras you can drag around. Every section is a small playground — move the cameras, break the assumptions, watch the geometry respond.
The problem: a camera flattens the world
Setup
A camera takes a 3D scene and squashes it onto a flat 2D sensor. That squashing throws away one whole dimension — depth. A point that's close and small looks exactly like a point that's far away and big, as long as they sit on the same ray from the camera's center. This is perspective projection, and it's the starting point for everything in this series: epipolar geometry, homographies, triangulation, and eventually reconstructing full 3D scenes from nothing but a handful of photos (structure from motion) or a moving robot's camera feed (SLAM).
The pinhole camera model
Foundations
Idealize a camera as a single point — the camera center C — through which every incoming light ray passes before hitting a flat image plane. A 3D point X in front of the camera projects to wherever the ray from C through X crosses that plane. Two steps turn a 3D world point into a 2D pixel:
1. World → camera. Express the point relative to the camera, using the camera's rotation R (which way it's facing) and position C:
2. Camera → pixel. Divide by depth (that's the "perspective" part — things farther away shrink) and scale into pixel units with the intrinsic matrix K:
v = fy · y_cam/z_cam + cy
Written as one matrix equation in homogeneous coordinates, with [R | −RC] called the camera's extrinsics:
fx, fy are the focal length in pixels along each axis, and (cx, cy) is the principal point — where the optical axis pierces the image, usually near the image center. K only depends on the camera's internals (lens, sensor); R and C only depend on where the camera is sitting in the world. That split matters — you'll reuse it in every part of this series.
Play with a camera
Interactive
Left: the scene in 3D — a cube and a camera (its little pyramid points at where it's looking). Right: exactly what that camera sees, computed live from the projection equation above. Drag to orbit the 3D view; use the sliders to move the camera around the cube and change its lens.
Drag to orbit the 3D scene, scroll to zoom.
What the camera sees — computed from x̃ = K[R|−RC]X̃.
Try dragging focal length up and down: a longer focal length is a "zoom lens" — it magnifies the cube but shrinks the field of view. Then try principal point offset: it shifts the whole image sideways without changing the camera's position at all — it's a property of the sensor, not the world.
The information a single camera throws away
Why one view isn't enough
Every point along a camera's viewing ray projects to the exact same pixel. A single image can't distinguish "small and close" from "big and far" — depth is genuinely lost. Drag the slider below: it slides a point back and forth along the camera's ray through a fixed pixel. Watch the 3D position change a lot, and the projected pixel not move at all.
Drag to orbit. The dashed line is the camera's ray through one fixed pixel; faint dots mark points at different depths along it.
The image plane: every point on that ray lands on the same pixel. The slider moves the 3D point; this image never changes.
Focal length, sensor size, and field of view
From millimeters to pixels
The sliders in Step 2 called fx, fy "focal length in pixels," but a real lens spec sheet talks in millimeters and a sensor spec talks in millimeters too. The bridge between them is simple trigonometry. A ray hitting the very edge of a sensor of width w (mm), a distance f (mm) behind the pinhole, makes an angle θ with the optical axis where tan(θ) = (w/2) / f — half the sensor width over the focal length. The full field of view is twice that:
To get from millimeters to the fx pixels this page's K actually uses, divide by the physical size of one pixel — the pixel pitch p = (sensor width in mm) / (image width in pixels):
Grounding it in a real number. A typical modern phone's main camera is quoted as "27mm equivalent" — a focal length normalized to a 36mm-wide full-frame sensor, so its FOV is 2·atan(36/(2·27)) = 2·atan(0.667) ≈ 67° horizontally, regardless of the phone's actual (much smaller) sensor. Suppose the real sensor is 4000 pixels wide with a 1.4µm pixel pitch — a physical width of 4000 · 0.0014mm ≈ 5.6mm. Solving the FOV formula backward for the real focal length gives f = (w/2) / tan(FOV/2) = 2.8 / tan(33.5°) ≈ 4.2mm, and in pixels that's fx = 4.2 / 0.0014 ≈ 3000px. So when a paper or a calibration file reports "fx ≈ 3000" for a 4000px-wide phone photo, that's a completely ordinary, slightly-wide-of-normal phone lens — not a special number. Compare that to Step 2's slider range (150–1200px on a 480px-wide canvas): a canvas that narrow needs proportionally smaller pixel focal lengths to show the same field of view, which is why fx=1200 there already looks like a telephoto zoom.
Side view: rays through the sensor edges define the FOV wedge. The same trigonometry, drawn to scale from the slider values.
Coordinate convention traps
The bug that isn't a bug
This series uses the OpenCV / computer-vision convention: from the camera's point of view, x points right, y points down (matching image row order — row 0 is the top), and z points forward, out along the optical axis, into the scene. Check the handedness: x × y = (1,0,0) × (0,1,0) = (0,0,1) = z — it's a perfectly ordinary right-handed frame, it just feels left-handed to people used to math-class axes because y is flipped from "up."
OpenGL and most classic graphics/robotics-visualization code instead use y up, x right, and z backward — the camera looks down its own −z axis, so anything visible has negative camera-space z. Also right-handed (x × y = z still holds), but a 180° rotation apart from the CV convention around the x-axis. COLMAP, a common SfM/MVS tool referenced later in this series, uses the same convention as this page (x-right, y-down, z-forward, world-to-camera X_cam = RX_world + t) — one less conversion if you're pulling data from it. Robotics body frames (ROS's REP-103) go a third way again: x-forward, y-left, z-up, meant for a vehicle chassis rather than a camera, so a robot's camera driver almost always publishes an explicit camera_link → base_link rotation to bridge the two worlds.
| Convention | x | y | z (camera looks along) | Handed |
|---|---|---|---|---|
| OpenCV / this series / COLMAP | right | down | +z (forward) | right |
| OpenGL / classic graphics | right | up | −z (backward) | right |
| ROS body frame (REP-103) | forward | left | z is up, not forward | right |
| Blender world | right | forward | z is up, camera looks along local −z | right |
If a rotation "looks weird," check these before assuming your math is wrong:
| Symptom | Usual cause |
|---|---|
| Image looks vertically flipped | Mixed a y-down (CV) source with a y-up (OpenGL) renderer without flipping v → H−v |
| Points that should be visible are "behind" the camera | Forward axis sign mismatch — CV convention is +z-forward, OpenGL is −z-forward |
| Everything rotated 180° | Row-vector vs column-vector convention mismatch: Rv vs vR transposes the rotation |
| Scene is mirror-imaged (correct up to a reflection) | Accidentally landed on a left-handed frame — usually one axis sign flipped when converting between conventions |
| Depth values are negative | Working in an OpenGL-style camera frame where visible points have z_cam < 0, not a bug in triangulation |
Lens distortion: the pinhole model's first lie
Brown–Conrady model
Real lenses aren't pinholes — they're curved glass, and that curvature bends light differently near the edges of the frame than at the center. The standard fix (the Brown–Conrady model) applies a correction in normalized camera coordinates (x, y) = (X_cam/Z_cam, Y_cam/Z_cam), before the linear K multiply — distortion is a property of the lens, not the sensor's pixel grid. Let r² = x² + y² (squared distance from the optical axis):
tangential (decentered lens): Δx = 2p₁xy + p₂(r²+2x²) Δy = p₁(r²+2y²) + 2p₂xy
x_distorted = xr + Δx y_distorted = yr + Δy
Positive k₁ pushes points away from center faster than linearly (pincushion); negative k₁ pulls them in (barrel — the classic wide-angle "fisheye bulge"). k₂, k₃ refine the curve further out; p₁, p₂ capture a lens that isn't perfectly centered over the sensor. Only after this correction does the pixel formula from Step 1 apply: x̃ = K·(x_distorted, y_distorted, 1).
x̃ = K[R|t]X̃, a strictly linear (projective) relationship. Distortion is not linear, and it breaks the one property projective transforms are supposed to preserve: straight 3D lines stop projecting to straight image lines. That's why every real pipeline undistorts the image (or just the matched point coordinates) as a preprocessing step, before any of the linear algebra in Parts 5–9 is allowed to touch the data.Undistorting means inverting the model above: given an observed distorted point (x_d, y_d), find the (x, y) that produced it. There's no closed form, so OpenCV's undistortPoints (and this demo) use a few steps of fixed-point iteration, starting from the distorted point itself as a first guess:
Play with it below: the grid is drawn as it would actually land on the sensor — perfectly straight world lines, forward-distorted by the sliders' k/p values. Toggle "recovered" to run the iteration above on those same distorted points and watch the straight grid come back.
Solid: distorted (what the sensor sees). Dashed, when toggled: recovered via iterative undistortion.
The same effect on a real scene: the cube projected with the sliders’ distortion. Straight 3D edges (gray dashed = pinhole truth) bend on the sensor — this is why Parts 6–9 require undistorted points.
Vanishing points and the horizon line
Where parallel lines meet — and what it buys you
Part 0 built the point at infinity for a 2D direction; the exact same idea exists one dimension up. A 3D direction d has an ideal point X̃∞ = (d, 0) — a homogeneous 4-vector with a zero last coordinate, representing "infinitely far away, along d." Project it through the ordinary camera equation:
The translation t drops out entirely — only vanishes because it's multiplied by the ideal point's zero last coordinate. So every 3D line running parallel to direction d, no matter where it sits in space, projects to a 2D line through the same fixed image point v = K·R·d — the vanishing point of that direction. This is literally "the image of the point at infinity," not a separate phenomenon: it's the same machinery from Part 0's x̃ ≅ ỹ and ideal points, just run through a camera.
A plane's directions all lie in a 2D subspace, so their vanishing points all lie on one image line — the image of that plane's line at infinity. For the ground plane, that's the horizon line: the vanishing points of any two non-parallel horizontal directions determine it completely, by exactly the same construction as Part 0's duality demo — l = v₁ × v₂.
The payoff — single-view metrology. If a scene has three mutually orthogonal directions (the corner of a room, a cube, most buildings), their three vanishing points constrain K itself. Each orthogonal pair (di, dj) gives one linear constraint on the image of the absolute conic ω = K⁻ᵀK⁻¹: viᵀ ω vj = 0 — vi and vj are said to be conjugate with respect to ω. Assuming zero skew and square pixels and a known principal point (cx,cy) (so K⁻¹v = ((u−cx)/f, (v−cy)/f, 1)), that constraint collapses to one clean equation for the focal length:
Recovering intrinsics (and, with all three pairs, the full K) from vanishing points alone, with no ruler and no checkerboard — one photo of a rectangular room is enough. Part 5 defines ω properly (it's the image of the absolute conic, a stand-in for K that ignores camera position) and shows where this constraint actually comes from.
Drag to orbit. Cube edges come in 3 parallel families along the world x/y/z axes.
Dashed: cube edges extended to their vanishing point. Solid horizontal line: the horizon (vpx × vpz).
The x-direction and z-direction vanishing points (both "horizontal" cube edge families) recover fx from nothing but their pixel positions via the formula above — watch recovered f track the true f slider as you orbit and re-zoom the camera.
Why a pinhole is an idealization
The assumptions everything else quietly relies on
Two more simplifications are baked into x̃ = K[R|t]X̃ that are worth naming honestly, even without an interactive demo attached.
Aperture and diffraction. A true point-sized pinhole lets through essentially no light, and worse, an infinitesimally small aperture diffracts light — the image actually gets blurrier as the hole shrinks past a certain point, not sharper. Real cameras trade the ideal point aperture for a finite lens opening, which lets in usable light but focuses only one distance perfectly; everything nearer or farther blurs into a small disk (the circle of confusion), which is exactly what "depth of field" is. The pinhole model this whole series uses is the f-number → ∞ limit of a real lens — everything in perfect focus at every depth, which no physical camera actually achieves.
Rolling shutter. Every equation in this series treats an image as one instantaneous snapshot: every pixel captured at the exact same instant, from the exact same camera pose. Most CMOS sensors don't work that way — they expose row by row, so the last row of a frame can be captured milliseconds after the first. For a static scene this is invisible; for a fast-moving camera (a car, a drone, a phone being waved around) or a fast-moving subject, different rows are effectively different camera poses, and treating the frame as one rigid snapshot introduces real geometric error. SLAM and structure-from-motion systems built for fast motion either use global-shutter cameras or explicitly model the per-row pose as a smooth trajectory (rolling-shutter bundle adjustment) rather than a single [R|t].
The general projective camera P
Putting it all in one 3×4 matrix
Collect everything from Step 1 into a single 3×4 matrix P = K[R|t]. Written out, P is exactly Part 0's "camera is a linear map P³ → P²" made concrete:
The camera center is the null space of P. Recall from Step 1 that t = −RC. Plug the camera center C's homogeneous coordinates C̃ = (C, 1) into P:
C is the one 3D point P sends to the zero vector — it has no well-defined image because it's not a direction from itself, it's the projection center itself. Solved the other way, C = −M⁻¹p₄ (a finite point, whenever M is invertible) — this is how you recover a camera's position directly from a raw P matrix, with no separate R, t needed.
Back-projection. Forward projection loses depth (Step 3); the reverse map — a pixel back out into a 3D ray — is exactly the set of points that P maps to that pixel. The ray through camera center C in direction M⁻¹x̃ works: parametrize X(λ) = C + λM⁻¹x̃ and apply M and add p₄:
— every point on that ray projects back to the same pixel, for any λ, which is the depth ambiguity from Step 3 stated as one line of algebra. (M's third row also gives the principal ray — the 3D direction the optical axis points, det(M)·m₃ — the sign fixing which way is "in front.")
Recovering K and R from a raw P: RQ decomposition. Given only the numbers in P (e.g. from a calibration routine that fit P directly with no K/R split), factor the leading 3×3 block M = K·R — upper-triangular K times orthogonal R. This is an RQ decomposition, not the more familiar QR: QR factors a matrix as (orthogonal)×(upper-triangular) — the wrong order for what's needed here, since K (upper-triangular) has to come first and the rotation R (orthogonal) second to match M = KR. In practice it's computed by reversing rows/columns, running ordinary QR, and reversing back — with a final sign fix-up so K's diagonal comes out positive (focal lengths shouldn't be negative) and det(R) = +1 (a proper rotation, not a reflection).
Finite vs. infinite cameras. Everything above assumes det(M) ≠ 0 — a finite camera with an ordinary 3D center. When det(M) = 0, M has no inverse and the "camera center" pushed out by solving PC̃=0 turns out to have a zero last coordinate — the center itself is a point at infinity. That's an affine (orthographic / weak-perspective) camera: parallel rays instead of rays converging on a finite point, no perspective foreshortening. It's the right model for a very distant or very long-lens camera, where the scene's own depth range is negligible next to the viewing distance and perspective effects are too small to matter.
Cheat sheet
Recap
| Piece | Symbol | Depends on | What it captures |
|---|---|---|---|
| Extrinsics | [R | −RC] | Where the camera sits | Position C and orientation R in the world |
| Intrinsics | K | The camera itself | Focal length, principal point (lens + sensor geometry) |
| Projection | x̃ = K[R|−RC]X̃ | Both | The full map from a 3D world point to a 2D pixel |
| What's lost | Depth z_cam | — | Every point on a ray from C maps to the same pixel |
| Field of view | FOV = 2·atan(w / 2f) | Lens + sensor | Relates focal length (mm) to sensor width and pixel-space fx |
| Lens distortion | xr = x(1+k₁r²+k₂r⁴+k₃r⁶) + tangential | Lens imperfection | Applied in normalized coords before K; must be undone before E/F/H are valid |
| General camera | P = K[R|t], 3×4 | Both | The camera as a linear map P³ → P²; x̃ = PX̃ |
| Camera center | C̃ = null(P) | — | PC̃ = 0; the one point with no well-defined image |
| Intrinsics/extrinsics from P | RQ decomposition | — | M = KR (upper-triangular × orthogonal, note the order vs. QR) |
| Vanishing point | v = K·R·d | Direction d | Image of the point at infinity X̃∞=(d,0) along d |