Beyond: SLAM, MVS & Learned Geometry
Every part of this series so far — from Part 4's projection equation through Part 16's full structure-from-motion pipeline — ends at the same kind of answer: a sparse point cloud, with 3D coordinates only for the handful of pixels that got matched as features. This closing part is a survey of what the field builds on top of that. First, dense multi-view stereo, which pushes from "a few thousand matched points" to "a depth estimate at nearly every pixel," and fuses many such depth maps into one solid surface. Then the newest shift: neural and feed-forward methods (NeRF, 3D Gaussian Splatting, DUSt3R-style models) that replace pieces of this classical pipeline with learned components — while, in every case that has actually stuck, still leaning on the same projection equation, known or estimated camera poses, and multi-view consistency this guide spent eleven parts deriving. Finally, SLAM, the real-time sibling of everything covered so far — the same tracking-and-mapping problem, but online, frame by frame, under a latency budget. The goal here is to place these on the map you've already built, not to rebuild any of them from scratch.
From sparse points to the rest of the field
Setup
Part 15 and Part 16 end with camera poses for every image and a triangulated 3D point for every feature that survived matching, verification, and bundle adjustment — typically a few thousand to a few hundred thousand points, clustered wherever the scene had enough texture to match reliably and empty everywhere else (blank walls, sky, anything low-contrast). That's a real reconstruction, but it's a skeleton, not a surface. Three things build on top of it, and it's worth being precise about how each one relates to everything already derived, because none of them throws it away:
Dense multi-view stereo: plane-sweep
Derivation
Sparse reconstruction only ever answers "where is this point" for points that were explicitly detected and matched as features (Part 7's corners/blobs, run through Part 7's RANSAC-verified matches). Dense multi-view stereo (MVS) asks a different question: given a reference image and one or more other images with already-known poses — the output of Part 16's SfM, treated here purely as an input — estimate a depth for (nearly) every pixel in the reference image, whether or not that pixel was ever part of a feature match.
The classical algorithm for this is plane-sweep stereo, and it's a direct extension of Part 8's plane-induced homography. Put the reference camera at the origin of its own coordinate frame (identity rotation, so its own optical axis is the local Z-axis), and consider a family of candidate fronto-parallel planes — planes perpendicular to that optical axis, at increasing depth Z. A plane at depth Z has normal n = (0,0,1) and distance d = Z, both expressed in the reference camera's own frame. For a second camera related to the reference by extrinsic pose X2 = R X1 + t, Part 8's plane-induced homography specializes to:
Hpix(Z) = K₂ H(Z) K₁⁻¹ (pixel coordinates — same K₂(⋅)K₁⁻¹ sandwich as every homography in Part 8)
(The sign in front of t nᵀ depends on exactly how the plane's offset d and normal direction are booked — Part 8 wrote H = R − t nᵀ/d for a plane measured the other way; same geometric object, opposite sign convention. Either is fine as long as it's applied consistently, which is all that matters for the demo below.)
This is the whole mechanism: for each candidate depth Z, Hpix(Z) tells you exactly where every reference pixel would land in the other image if the scene were really a flat plane sitting at depth Z. Warp the other image backward through that homography so it lines up with the reference frame, then compare colors: wherever the warped pixel actually agrees with the reference pixel — high photo-consistency, scored with the same SAD/SSD/NCC window metrics Part 9 used for disparity — that candidate Z is probably close to the true depth at that pixel. Sweep Z across the scene's depth range, and independently for every pixel, keep whichever Z scored best. That per-pixel best-Z map is the dense depth map.
H(Z) assumes the whole scene is the flat plane at depth Z, which is false almost everywhere. But photo-consistency is scored independently per pixel, so that assumption only needs to be locally true — true for the small patch of pixels that actually sit near depth Z — for the score to peak correctly there. Pixels belonging to a different surface at a different depth simply score low at this Z and high at their own Z instead, once the sweep gets there.PatchMatch MVS. Sweeping every candidate depth at every pixel is correct but wasteful: most of a cost volume's entries are far from any real surface and never win. Practical systems (COLMAP's dense stage, and PatchMatch-based MVS generally) borrow the PatchMatch algorithm from image editing: instead of exhaustively testing all depths everywhere, start from a random depth guess per pixel, then repeatedly (1) propagate — try each pixel's neighbors' current best depth, since real surfaces are usually locally smooth and a good guess next door is probably a good guess here too, and (2) randomly perturb the current best guess by a shrinking amount, in case propagation alone gets stuck in the wrong neighborhood. A handful of these propagate-and-perturb rounds converges to essentially the same depth map as brute-force sweeping, for a small fraction of the cost, which is exactly why it's the practical default rather than a curiosity.
Depth fusion: from many depth maps to one surface
Merging
MVS is run once per reference view, so a scene photographed from fifty images yields fifty depth maps — each one noisy, each one seeing a different part of the scene, and none of them individually consistent with the others (the same physical surface, estimated from two different reference views, rarely comes back at exactly the same depth). Turning that pile of overlapping, disagreeing depth maps into one coherent 3D model is depth fusion, and there are two standard routes.
TSDF fusion. Lay a regular 3D voxel grid over the scene. For every voxel, and every depth map that "sees" it (i.e. the voxel projects into that camera's field of view), compute the signed distance from the voxel to the surface implied by that depth map along the corresponding camera ray — positive in front of the surface, negative behind it, and truncated to a small band near zero so that voxels far from any observed surface don't get polluted by one noisy reading (hence Truncated Signed Distance Function). Average that signed distance into the voxel across every depth map that observes it. The result is a smooth scalar field over the whole volume, and the surface is defined implicitly as wherever that field crosses zero — extractable as an explicit mesh with a standard algorithm like marching cubes. Averaging many noisy per-view estimates into each voxel is what makes TSDF fusion robust to the disagreement between individual depth maps: no single view's error dominates.
Poisson surface reconstruction. The point-based alternative: instead of a depth-map-per-voxel average, start from an oriented point cloud — the fused 3D points from MVS (or even directly from a dense point set), each carrying an estimated surface normal. Poisson reconstruction fits a smooth implicit function whose gradient best matches those normals everywhere, by solving a Poisson equation (a specific linear PDE — the same "Poisson" as in Poisson's equation from physics) over the volume; the output surface is again the zero level set of that function, extracted the same way. It tends to produce smoother, more watertight meshes than raw TSDF fusion, at the cost of needing decent normal estimates going in.
Cross-section (x horizontal, depth vertical). Dashed lines: the two individual depth maps’ surfaces; black solid: the TSDF zero crossing extracted from the fused field.
Play: plane-sweep MVS on a synthetic scene
Interactive
Hpix(Z), warps the other image through it, and scores real windowed SAD photo-consistency per pixel. Nothing here is scripted or pre-baked; this is not a production MVS/PatchMatch implementation (no multi-view robustness, no confidence filtering, no real image content) but the mechanism it's illustrating — plane-sweep photo-consistency — is computed exactly as described above.Reference view
Other view
Photo-consistency at current Z
Heatmap: bright = the warped "other view" agrees well with the reference view at this candidate depth (low SAD over a small window); dark = it doesn't.
Average photo-consistency across the whole image, plotted against swept depth Z. Dashed markers = the two true surface depths (box top, floor); the curve should peak right at them.
The end product of the sweep: per-pixel best-Z rendered as brightness (bright = nearer, dark = farther). Two distinct depth layers visible: the raised box and the floor around it — this is a dense depth map, no features matched anywhere.
NeRF, 3D Gaussian Splatting, feed-forward geometry
Same geometry, new representation
The methods that have generated the most recent excitement in this field are not a replacement for anything above — they're new answers to "what does the scene look like," built on the same projection equation, the same camera poses, and the same multi-view-consistency principle ("many views of the same real scene must agree") that every earlier part of this series derived. What changes is how the scene is stored, and in one case, how the whole pipeline is solved.
NeRF (Neural Radiance Fields). Instead of an explicit point cloud, mesh, or voxel grid, represent the scene as a neural network: a function that takes a 3D position and a viewing direction and outputs a color and a volume density. To render a pixel, cast the exact same ray Part 4's pinhole model back-projects for that pixel, sample the network at points along it, and integrate — weighting each sampled color by its density and by how much light survives to reach it (an occluding surface earlier along the ray attenuates everything behind it):
Training minimizes the difference between rendered and observed pixels across every training photo — but critically, the camera pose for every one of those photos comes from an ordinary Part-16-style SfM run beforehand, fed in as a fixed input. NeRF does not solve for camera poses (variants that jointly refine them exist, but the baseline algorithm assumes SfM already happened); it only replaces what the scene is made of once poses are known.
3D Gaussian Splatting. Keep the explicit-representation idea, but drop the network: represent the scene as a set of 3D Gaussians, each with a position, a covariance (shape/orientation), a color, and an opacity. Rendering doesn't march rays through a network — it rasterizes the Gaussians directly, projecting each one onto the image plane with the same projection math as Part 4 and compositing the overlapping splats front-to-back, GPU-rasterizer style. That's why it renders dramatically faster than NeRF: rasterizing a fixed, explicit set of primitives is cheap and embarrassingly parallel, where NeRF has to evaluate a neural network at many sample points along every single ray, for every pixel, every frame. What doesn't change is the training loop: render from a known pose, compare to the real photo, backpropagate the pixel error into the Gaussians' parameters — the same "render, compare, backpropagate" cycle as NeRF, just through a rasterizer instead of a ray-marcher, and again starting from SfM poses.
Feed-forward geometry (DUSt3R / MASt3R / VGGT-family models). The newest shift touches the solve, not just the representation. Part 16's classical pipeline is a multi-stage optimization: detect and match features, verify matches geometrically (Part 7's RANSAC), then solve incrementally or globally for poses and structure (Part 16). Feed-forward models replace that whole multi-stage pipeline with a single neural network that takes in a pair or a small set of images and, in one forward pass, directly regresses dense 3D point correspondences (and often camera poses) — no explicit matching stage, no iterative bundle adjustment. The network is trained end-to-end on large datasets to implicitly learn what the classical pipeline computes explicitly.
Play: the same scene, four representations
Interactive
Same 3D samples, same camera, four different renderers.
SLAM & visual-inertial odometry: the online sibling
From batch to real time
Every part of this series so far has been offline / batch: given a fixed set of images collected in advance, reconstruct once, taking as long as needed. SLAM (simultaneous localization and mapping) solves the fundamental same problem — track where the camera is (Part 11's PnP-style pose-from-known-3D-points) while building and extending a 3D map of the scene (Part 10's triangulation) — but online, one frame at a time, as a robot or headset moves through the world, under a strict latency budget: the answer for this frame has to be ready before the next one arrives.
That latency constraint is what reshapes the architecture. A batch SfM run can afford to jointly optimize every pose and every point together (Part 15's full bundle adjustment) because it only has to finish once. A SLAM system can't wait that long every frame, so it splits the work:
Front end. Runs every single frame, fast: track features from the previous frame (or match against the existing local map), and estimate the current camera pose from those correspondences — essentially Part 11's PnP problem, solved against 3D points the map already has, repeated 20-60 times a second.
Back end. Runs slower, usually on a separate background thread, and doesn't touch every frame: a windowed (or "local") bundle adjustment — the same nonlinear least-squares machinery as Part 15, but restricted to a sliding window of recent keyframes rather than the whole trajectory — periodically refines poses and map points to keep small per-frame errors from accumulating unbounded. The front end reads the back end's latest refined estimate but never waits on it.
Loop closure. Incremental SfM (Part 16) already flagged that chaining relative poses frame-by-frame accumulates drift with no way to self-correct. SLAM's fix is to actively look for it: an image-retrieval / place-recognition step continuously checks whether the current view matches somewhere the system has already mapped. When it finds a match — "this is the same corner I saw ten minutes ago" — that becomes a new constraint between two keyframes that are far apart in time but geometrically close, fed into a pose-graph optimization: a lighter-weight cousin of full bundle adjustment that optimizes only the relative poses between keyframes (not the full 3D map alongside them), cheap enough to run globally and snap the accumulated drift back into consistency the moment a loop closes.
Red = drifted odometry-only trajectory; teal = current pose-graph estimate as the loop-closure edge’s weight grows. Full derivation of pose-graph optimization: Nonlinear Optimization: pose graphs.
Cheat sheet
Recap
| Object | What it does | Notes |
|---|---|---|
| Plane-sweep MVS | Sweep fronto-parallel depth planes; warp other view via Hpix(Z)=K₂(R+tnᵀ/Z)K₁⁻¹; keep the Z with best photo-consistency, per pixel | Part 8's plane homography, specialized to n = the reference camera's own axis |
| PatchMatch MVS | Propagate neighbors' good depth guesses + random perturbation, instead of exhaustive sweep | Converges far faster; COLMAP's dense stage uses this |
| TSDF fusion | Voxel grid of signed distance to nearest surface, averaged across all depth maps | Surface = zero crossing; extract with marching cubes |
| Poisson reconstruction | Fits a smooth implicit function to an oriented point cloud (points + normals) | Zero level set = mesh; smoother than raw TSDF, needs good normals |
| NeRF | Network(position, direction) → color, density; volume-rendered along the Part 4 ray | Same geometry, new representation — poses still come from ordinary SfM (Part 16) |
| 3D Gaussian Splatting | Explicit, differentiable 3D Gaussians, rasterized with Part 4's projection math | Same geometry, new representation — rasterizing beats per-ray network evaluation |
| Feed-forward geometry | One network regresses dense point maps / poses in a single forward pass | Same geometry, new solve — replaces Part 16's multi-stage pipeline; newest, still maturing |
| SLAM front end | Per-frame feature tracking + PnP-style pose estimate against the existing map | Runs every frame, has to be fast |
| SLAM back end | Windowed bundle adjustment (Part 15) over recent keyframes | Background thread; front end never waits on it |
| Loop closure | Place recognition → new constraint → pose-graph optimization | Lighter cousin of full BA; fixes the drift Part 16 flagged |
| Visual-inertial | Fuses high-rate IMU data with vision | Bridges fast motion between frames; fixes the scale ambiguity from Parts 10 & 16 |