Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

0

From sparse points to the rest of the field

Setup

Part 15 and Part 16 end with camera poses for every image and a triangulated 3D point for every feature that survived matching, verification, and bundle adjustment — typically a few thousand to a few hundred thousand points, clustered wherever the scene had enough texture to match reliably and empty everywhere else (blank walls, sky, anything low-contrast). That's a real reconstruction, but it's a skeleton, not a surface. Three things build on top of it, and it's worth being precise about how each one relates to everything already derived, because none of them throws it away:

🧭 Roadmap: dense multi-view stereo keeps the batch/offline setting and the known poses from SfM, but changes the question — instead of "where is this matched feature," it asks "what is the depth at every pixel." Neural and feed-forward methods keep the same projection equation and multi-view consistency principle, but change the representation of the scene (a network, a set of Gaussians) or the solve (one forward pass instead of an iterative pipeline). SLAM keeps the same tracking-and-mapping problem, but changes the timing — online and frame-by-frame instead of batch. Same geometry throughout; what changes is what's being asked, how the scene is stored, or when the answer is needed.
1

Dense multi-view stereo: plane-sweep

Derivation

Sparse reconstruction only ever answers "where is this point" for points that were explicitly detected and matched as features (Part 7's corners/blobs, run through Part 7's RANSAC-verified matches). Dense multi-view stereo (MVS) asks a different question: given a reference image and one or more other images with already-known poses — the output of Part 16's SfM, treated here purely as an input — estimate a depth for (nearly) every pixel in the reference image, whether or not that pixel was ever part of a feature match.

The classical algorithm for this is plane-sweep stereo, and it's a direct extension of Part 8's plane-induced homography. Put the reference camera at the origin of its own coordinate frame (identity rotation, so its own optical axis is the local Z-axis), and consider a family of candidate fronto-parallel planes — planes perpendicular to that optical axis, at increasing depth Z. A plane at depth Z has normal n = (0,0,1) and distance d = Z, both expressed in the reference camera's own frame. For a second camera related to the reference by extrinsic pose X2 = R X1 + t, Part 8's plane-induced homography specializes to:

H(Z) = R + t nᵀ / Z   (calibrated coordinates; n=(0,0,1)ᵀ, the reference camera's own forward axis)
Hpix(Z) = K₂ H(Z) K₁⁻¹   (pixel coordinates — same K₂(⋅)K₁⁻¹ sandwich as every homography in Part 8)

(The sign in front of t nᵀ depends on exactly how the plane's offset d and normal direction are booked — Part 8 wrote H = R − t nᵀ/d for a plane measured the other way; same geometric object, opposite sign convention. Either is fine as long as it's applied consistently, which is all that matters for the demo below.)

This is the whole mechanism: for each candidate depth Z, Hpix(Z) tells you exactly where every reference pixel would land in the other image if the scene were really a flat plane sitting at depth Z. Warp the other image backward through that homography so it lines up with the reference frame, then compare colors: wherever the warped pixel actually agrees with the reference pixel — high photo-consistency, scored with the same SAD/SSD/NCC window metrics Part 9 used for disparity — that candidate Z is probably close to the true depth at that pixel. Sweep Z across the scene's depth range, and independently for every pixel, keep whichever Z scored best. That per-pixel best-Z map is the dense depth map.

⚠️ Why this only works pixel-by-pixel, not globally: a single homography H(Z) assumes the whole scene is the flat plane at depth Z, which is false almost everywhere. But photo-consistency is scored independently per pixel, so that assumption only needs to be locally true — true for the small patch of pixels that actually sit near depth Z — for the score to peak correctly there. Pixels belonging to a different surface at a different depth simply score low at this Z and high at their own Z instead, once the sweep gets there.

PatchMatch MVS. Sweeping every candidate depth at every pixel is correct but wasteful: most of a cost volume's entries are far from any real surface and never win. Practical systems (COLMAP's dense stage, and PatchMatch-based MVS generally) borrow the PatchMatch algorithm from image editing: instead of exhaustively testing all depths everywhere, start from a random depth guess per pixel, then repeatedly (1) propagate — try each pixel's neighbors' current best depth, since real surfaces are usually locally smooth and a good guess next door is probably a good guess here too, and (2) randomly perturb the current best guess by a shrinking amount, in case propagation alone gets stuck in the wrong neighborhood. A handful of these propagate-and-perturb rounds converges to essentially the same depth map as brute-force sweeping, for a small fraction of the cost, which is exactly why it's the practical default rather than a curiosity.

2

Depth fusion: from many depth maps to one surface

Merging

MVS is run once per reference view, so a scene photographed from fifty images yields fifty depth maps — each one noisy, each one seeing a different part of the scene, and none of them individually consistent with the others (the same physical surface, estimated from two different reference views, rarely comes back at exactly the same depth). Turning that pile of overlapping, disagreeing depth maps into one coherent 3D model is depth fusion, and there are two standard routes.

TSDF fusion. Lay a regular 3D voxel grid over the scene. For every voxel, and every depth map that "sees" it (i.e. the voxel projects into that camera's field of view), compute the signed distance from the voxel to the surface implied by that depth map along the corresponding camera ray — positive in front of the surface, negative behind it, and truncated to a small band near zero so that voxels far from any observed surface don't get polluted by one noisy reading (hence Truncated Signed Distance Function). Average that signed distance into the voxel across every depth map that observes it. The result is a smooth scalar field over the whole volume, and the surface is defined implicitly as wherever that field crosses zero — extractable as an explicit mesh with a standard algorithm like marching cubes. Averaging many noisy per-view estimates into each voxel is what makes TSDF fusion robust to the disagreement between individual depth maps: no single view's error dominates.

Poisson surface reconstruction. The point-based alternative: instead of a depth-map-per-voxel average, start from an oriented point cloud — the fused 3D points from MVS (or even directly from a dense point set), each carrying an estimated surface normal. Poisson reconstruction fits a smooth implicit function whose gradient best matches those normals everywhere, by solving a Poisson equation (a specific linear PDE — the same "Poisson" as in Poisson's equation from physics) over the volume; the output surface is again the zero level set of that function, extracted the same way. It tends to produce smoother, more watertight meshes than raw TSDF fusion, at the cost of needing decent normal estimates going in.

🎯 Learning goal: TSDF fusion, in cross-section. Two noisy depth maps of the same floor-and-box scene vote on every voxel; each casts a truncated signed distance (red = in front of surface, blue = behind, white = zero). Slide the truncation distance and noise level, and watch the averaged zero crossing (black) sit where the true surface is — between the two maps' noisy disagreements, which is the whole trick. The truncation band is why voxels far from any surface never corrupt the field.

Cross-section (x horizontal, depth vertical). Dashed lines: the two individual depth maps’ surfaces; black solid: the TSDF zero crossing extracted from the fused field.

3

Play: plane-sweep MVS on a synthetic scene

Interactive

🎯 Learning goal: sweep the candidate depth plane and watch the live photo-consistency heatmap light up — dim almost everywhere, then bright exactly where the swept depth matches a real surface: first the near, raised patch, then the far, textured backdrop.
⚠️ Honest framing: the two small images below are ray-cast from a synthetic scene — a textured "floor" backdrop and a raised textured "box," both procedurally checkered so there's real pixel content to match — through two virtual pinhole cameras with exactly known poses, chosen nearly overhead so both surfaces are close to fronto-parallel to the reference camera. That last part is a deliberate simplification: it's what makes the demo's peaks land cleanly on a single depth rather than smearing across a range, the way an oblique real-world ground plane would. Everything else is real, live computation — for every value of Z on the slider, the code below actually builds Hpix(Z), warps the other image through it, and scores real windowed SAD photo-consistency per pixel. Nothing here is scripted or pre-baked; this is not a production MVS/PatchMatch implementation (no multi-view robustness, no confidence filtering, no real image content) but the mechanism it's illustrating — plane-sweep photo-consistency — is computed exactly as described above.

Reference view

Other view

Photo-consistency at current Z

Heatmap: bright = the warped "other view" agrees well with the reference view at this candidate depth (low SAD over a small window); dark = it doesn't.

Average photo-consistency across the whole image, plotted against swept depth Z. Dashed markers = the two true surface depths (box top, floor); the curve should peak right at them.

The end product of the sweep: per-pixel best-Z rendered as brightness (bright = nearer, dark = farther). Two distinct depth layers visible: the raised box and the floor around it — this is a dense depth map, no features matched anywhere.

4

NeRF, 3D Gaussian Splatting, feed-forward geometry

Same geometry, new representation

The methods that have generated the most recent excitement in this field are not a replacement for anything above — they're new answers to "what does the scene look like," built on the same projection equation, the same camera poses, and the same multi-view-consistency principle ("many views of the same real scene must agree") that every earlier part of this series derived. What changes is how the scene is stored, and in one case, how the whole pipeline is solved.

NeRF (Neural Radiance Fields). Instead of an explicit point cloud, mesh, or voxel grid, represent the scene as a neural network: a function that takes a 3D position and a viewing direction and outputs a color and a volume density. To render a pixel, cast the exact same ray Part 4's pinhole model back-projects for that pixel, sample the network at points along it, and integrate — weighting each sampled color by its density and by how much light survives to reach it (an occluding surface earlier along the ray attenuates everything behind it):

C(ray) = ∫ T(s)·σ(s)·c(s) ds    T(s) = exp( −∫₀ₖ σ(u) du )  (accumulated transmittance to depth s)

Training minimizes the difference between rendered and observed pixels across every training photo — but critically, the camera pose for every one of those photos comes from an ordinary Part-16-style SfM run beforehand, fed in as a fixed input. NeRF does not solve for camera poses (variants that jointly refine them exist, but the baseline algorithm assumes SfM already happened); it only replaces what the scene is made of once poses are known.

3D Gaussian Splatting. Keep the explicit-representation idea, but drop the network: represent the scene as a set of 3D Gaussians, each with a position, a covariance (shape/orientation), a color, and an opacity. Rendering doesn't march rays through a network — it rasterizes the Gaussians directly, projecting each one onto the image plane with the same projection math as Part 4 and compositing the overlapping splats front-to-back, GPU-rasterizer style. That's why it renders dramatically faster than NeRF: rasterizing a fixed, explicit set of primitives is cheap and embarrassingly parallel, where NeRF has to evaluate a neural network at many sample points along every single ray, for every pixel, every frame. What doesn't change is the training loop: render from a known pose, compare to the real photo, backpropagate the pixel error into the Gaussians' parameters — the same "render, compare, backpropagate" cycle as NeRF, just through a rasterizer instead of a ray-marcher, and again starting from SfM poses.

Feed-forward geometry (DUSt3R / MASt3R / VGGT-family models). The newest shift touches the solve, not just the representation. Part 16's classical pipeline is a multi-stage optimization: detect and match features, verify matches geometrically (Part 7's RANSAC), then solve incrementally or globally for poses and structure (Part 16). Feed-forward models replace that whole multi-stage pipeline with a single neural network that takes in a pair or a small set of images and, in one forward pass, directly regresses dense 3D point correspondences (and often camera poses) — no explicit matching stage, no iterative bundle adjustment. The network is trained end-to-end on large datasets to implicitly learn what the classical pipeline computes explicitly.

⚠️ How new this is: these models are very recent, and the classical pipeline this whole series has built remains the backbone for high-precision, large-scale, or safety-critical reconstruction — its failure modes (drift, degenerate configurations, RANSAC's own limits from Part 7) are well characterized after two decades of use, which matters as much as raw accuracy. Feed-forward methods are currently strongest where the classical pipeline struggles or is simply too slow to bother with: fast turnaround, sparse or few-view inputs, and scenes with correspondence that's hard to match by hand-engineered features (texture-less regions, wide baselines). Which approach "wins" for a given job is still very much an open, moving target.
6

SLAM & visual-inertial odometry: the online sibling

From batch to real time

Every part of this series so far has been offline / batch: given a fixed set of images collected in advance, reconstruct once, taking as long as needed. SLAM (simultaneous localization and mapping) solves the fundamental same problem — track where the camera is (Part 11's PnP-style pose-from-known-3D-points) while building and extending a 3D map of the scene (Part 10's triangulation) — but online, one frame at a time, as a robot or headset moves through the world, under a strict latency budget: the answer for this frame has to be ready before the next one arrives.

That latency constraint is what reshapes the architecture. A batch SfM run can afford to jointly optimize every pose and every point together (Part 15's full bundle adjustment) because it only has to finish once. A SLAM system can't wait that long every frame, so it splits the work:

Front end. Runs every single frame, fast: track features from the previous frame (or match against the existing local map), and estimate the current camera pose from those correspondences — essentially Part 11's PnP problem, solved against 3D points the map already has, repeated 20-60 times a second.

Back end. Runs slower, usually on a separate background thread, and doesn't touch every frame: a windowed (or "local") bundle adjustment — the same nonlinear least-squares machinery as Part 15, but restricted to a sliding window of recent keyframes rather than the whole trajectory — periodically refines poses and map points to keep small per-frame errors from accumulating unbounded. The front end reads the back end's latest refined estimate but never waits on it.

Loop closure. Incremental SfM (Part 16) already flagged that chaining relative poses frame-by-frame accumulates drift with no way to self-correct. SLAM's fix is to actively look for it: an image-retrieval / place-recognition step continuously checks whether the current view matches somewhere the system has already mapped. When it finds a match — "this is the same corner I saw ten minutes ago" — that becomes a new constraint between two keyframes that are far apart in time but geometrically close, fed into a pose-graph optimization: a lighter-weight cousin of full bundle adjustment that optimizes only the relative poses between keyframes (not the full 3D map alongside them), cheap enough to run globally and snap the accumulated drift back into consistency the moment a loop closes.

⚠️ Visual-inertial (VIO/VI-SLAM): fusing an IMU's high-rate accelerometer and gyroscope measurements with the vision pipeline solves two separate problems at once. First, bridging fast motion between camera frames — a whip-pan or a bump can move the camera enough between two frames that visual tracking alone loses the match, while the IMU samples at hundreds of Hz and can carry the pose estimate through the gap. Second, and just as important: pure vision from a single moving camera can only ever recover structure and motion up to an unknown overall scale — the exact ambiguity Part 10 and Part 16 both flagged (a scene and its poses look identical whether it's a tabletop model or the real building, twice the size, twice as far away). An accelerometer measures physical acceleration in real units (m/s²), which pure vision has no equivalent of — integrating it twice gives an absolute physical distance reference that resolves the scale ambiguity outright.
🎯 Learning goal: the pose graph, the data structure behind loop closure. Camera nodes along a walked loop; odometry edges (thin) accumulated drift; one loop-closure edge (thick teal) added when place recognition fires. Slide the closure weight from 0 to 1 and watch the drift redistribute around the whole loop instead of piling up at the seam — the same relaxation the Part 16 demo ran, driven by a single slider.

Red = drifted odometry-only trajectory; teal = current pose-graph estimate as the loop-closure edge’s weight grows. Full derivation of pose-graph optimization: Nonlinear Optimization: pose graphs.

✓

Cheat sheet

Recap

ObjectWhat it doesNotes
Plane-sweep MVSSweep fronto-parallel depth planes; warp other view via Hpix(Z)=K₂(R+tnᵀ/Z)K₁⁻¹; keep the Z with best photo-consistency, per pixelPart 8's plane homography, specialized to n = the reference camera's own axis
PatchMatch MVSPropagate neighbors' good depth guesses + random perturbation, instead of exhaustive sweepConverges far faster; COLMAP's dense stage uses this
TSDF fusionVoxel grid of signed distance to nearest surface, averaged across all depth mapsSurface = zero crossing; extract with marching cubes
Poisson reconstructionFits a smooth implicit function to an oriented point cloud (points + normals)Zero level set = mesh; smoother than raw TSDF, needs good normals
NeRFNetwork(position, direction) → color, density; volume-rendered along the Part 4 raySame geometry, new representation — poses still come from ordinary SfM (Part 16)
3D Gaussian SplattingExplicit, differentiable 3D Gaussians, rasterized with Part 4's projection mathSame geometry, new representation — rasterizing beats per-ray network evaluation
Feed-forward geometryOne network regresses dense point maps / poses in a single forward passSame geometry, new solve — replaces Part 16's multi-stage pipeline; newest, still maturing
SLAM front endPer-frame feature tracking + PnP-style pose estimate against the existing mapRuns every frame, has to be fast
SLAM back endWindowed bundle adjustment (Part 15) over recent keyframesBackground thread; front end never waits on it
Loop closurePlace recognition → new constraint → pose-graph optimizationLighter cousin of full BA; fixes the drift Part 16 flagged
Visual-inertialFuses high-rate IMU data with visionBridges fast motion between frames; fixes the scale ambiguity from Parts 10 & 16
That's the series, start to finish — projective geometry to the frontier. Jump back to the start, or revisit any part from the nav above. ↺ Back to Part 0: projective geometry