Multi-View Geometry, Interactively
A step-by-step guide to the geometry that turns flat images into 3D structure, built around one running scene - a cube seen by one or more draggable pinhole cameras. Every part is interactive; nothing is a static diagram.
Parts build on each other, but each one stands alone. If you already know homogeneous coordinates, start at Part 4; if you are here for a specific estimator, jump straight to it. The structure-from-motion parts assemble everything into a working reconstruction pipeline, and lean on the nonlinear optimization guide for the solver underneath.
The parts
The projective plane, homogeneous coordinates, the congruence symbol, and the point/line duality every later part depends on.
A conic as the quadratic form x̃ᵀCx̃=0, how a projective map transforms it, and the absolute conic whose image ω=K⁻ᵀK⁻¹ is exactly what calibration recovers.
Euclidean ⊂ similarity ⊂ affine ⊂ projective: the degrees of freedom each group adds, the quantities each destroys, and cross-ratio as the invariant that survives.
Points and planes in P³ with their dual incidence rule, the plane at infinity, and Plücker coordinates (d, m) for lines together with the d·m = 0 constraint.
Perspective projection, focal length and field of view, coordinate-convention traps, lens distortion, and the general projective camera P.
How you actually get K: Zhang's method, checkerboard capture, the image of the absolute conic, skew, and what a good reprojection error looks like.
The epipolar constraint, essential and fundamental matrices, and why matching points between two images collapses to a 1D search along a line.
Outlier rejection with RANSAC and its descendants (MSAC, LO-RANSAC, MAGSAC), minimal samples, the iteration-count formula, and Sampson vs. algebraic error.
The planar homography, the pure-rotation special case, DLT estimation, and why it silently breaks once points leave the plane.
Rectifying an image pair so epipolar lines become horizontal scanlines, block matching and SGBM, disparity maps, and Z = fB/d.
Recovering relative camera pose from the essential matrix, and triangulating 3D points from two known views, linear vs. nonlinear.
Recovering absolute camera pose from 2D-3D correspondences: DLT resection, P3P, EPnP, and PnP+RANSAC, the way every new camera enters a reconstruction.
Why a third camera view is fully predictable from the first two, and the trifocal tensor that captures three-view geometry directly.
Why a calibrated pair needs only five correspondences: the essential matrix's equal-singular-value and rank-2 constraints, Nister's solver, and how it beats the eight-point fit on difficult scenes.
Under weak perspective the tracked features form a rank-3 measurement matrix; one SVD factors it into camera motion and 3D shape, with the affine gauge fixed by rotation orthonormality.
The reprojection-error cost over every camera and every point, why it connects straight back to Gauss-Newton and Levenberg-Marquardt, and a toy SfM demo you run yourself.
Incremental structure-from-motion (COLMAP-style) versus global SfM, self-calibration, and the stratified projective to affine to metric upgrade.
Recovering the image of the absolute conic from images alone: Kruppa's equations, the absolute dual quadric, and the linear solve that upgrades a projective reconstruction to a metric one.
Planar scenes, pure rotation, critical surfaces and near-degenerate baselines: the geometric coincidences that break the estimators, and how pipelines detect them.
Where the field goes next: SLAM, dense multi-view stereo, and the learned successors (NeRF, 3D Gaussian Splatting, DUSt3R/VGGT), framed as the same geometry in new representations.