Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

0

An estimator is only as good as the geometry feeding it

Motivation

Look back at the shape of every estimation step so far. The eight-point algorithm builds a matrix A from correspondences and takes its smallest singular vector. Homography DLT does the same with a different A. Triangulation solves a small linear system; PnP-RANSAC repeatedly solves a minimal one. In each case the answer is the null space of a matrix — and a null space is only a single, stable direction when the geometric configuration that built the matrix is generic.

A degenerate configuration is a specific, nameable coincidence in the scene and camera placement where that stops being true. There are exactly two ways it can go wrong. Either the constraint matrix loses rank — a second direction also nearly satisfies every row, so the solver's choice of "the" answer becomes arbitrary and noise-dominated. Or two genuinely different 3D explanations fit the same images equally well — a two-fold or continuous ambiguity that no amount of residual minimization can resolve, because both explanations have the same residual.

The important practical fact is that these cases do not announce themselves. A degenerate fit still returns an F, still gives a reprojection error near zero on the (degenerate) input, and still looks like a successful reconstruction. What distinguishes a robust pipeline from a naive one is that it checks for the coincidences this page describes, and falls back to a different model when it finds one.

2

Detection and model selection

Foundations

Because a degenerate fit reports a small residual, residual alone is useless as a detector. Robust pipelines instead ask three separate questions: which model is simpler that fits just as well, how well-determined is the constraint matrix, and does the geometry itself have enough parallax to support the answer.

Model selection: comparing H against F

The cleanest test for a planar scene is to fit both a homography (8 DOF) and a fundamental matrix (7 DOF) to the same correspondences and compare. A penalized score does the judging. The classic one is GRIC, the Geometric Robust Information Criterion: a sum of robust squared residuals plus a per-degree-of-freedom penalty λ·d·n (with n the number of points). Its whole point is that a more complex model must earn its extra parameters by explaining the data clearly better. If H is nearly as good as F, GRIC prefers H — which is exactly the correct call on a plane, and it also catches pure rotation, where the motion is a homography. A related, cheaper heuristic is simply to compare RANSAC inlier counts: when the homography model reaches almost the same inlier count as F, suspect planarity.

Conditioning of the constraint matrix

For a well-posed fit, the smallest singular value of A is clearly separated from the next one: the null space is a single, unambiguous direction. For a rank-deficient configuration the two smallest values crowd together, so the "smallest singular vector" is a poorly determined mixture of two candidates. The ratio σ₂/σ₁ (second-smallest over smallest) therefore reads large and healthy on a general scene and near one as the configuration degenerates. This is the single most useful scalar to watch, and it is reported live in the demo below.

A cheap geometric gate: baseline and triangulation angle

Before trusting any per-point depth, measure the angle between the two back-projected rays that meet at the point (equivalently, the parallax). Healthy two-view geometry has a healthy spread of these angles; pure rotation collapses them to zero, and a small baseline collapses them to a fraction of a degree. A median-angle threshold is a fast, model-free way to discard a bad pair or mark a region as depth-unreliable — and it is exactly the quantity the depth-uncertainty formula ΔZ ≈ σ·Z²/(fB) is punishing.

Degeneracy checks on the essential matrix

Once the cameras are calibrated, F is promoted to E = Kᵀ F K, and a valid essential matrix has a very specific spectrum: two equal nonzero singular values and a third that is zero, σ₁ = σ₂ > σ₃ = 0. The shape of that spectrum is itself a diagnostic. If the two nonzero values come out nearly equal, the matrix is genuinely an essential matrix; if one of them collapses toward zero (so the matrix is closer to rank 1), the estimated geometry is sitting in or near a degenerate configuration and should not be trusted as calibrated motion.

🎯 Why modern SfM fits multiple models per pair. Incremental structure-from-motion (the COLMAP-style pipeline) does not assume the general case and hope. For every image pair it initializes both a homography and a fundamental matrix, runs RANSAC for each, and picks the winner by inlier count and GRIC-style scoring. Pairs that choose the homography are flagged so the pipeline does not try to triangulate them; frames whose motion is dominated by rotation are handled without pretending to recover depth. Model selection is not a refinement — it is the difference between a pipeline that quietly fails on a rotating phone video and one that does not.
3

Play: switch on a degeneracy and watch the estimator break

Interactive

🎯 Learning goal: the same eight-point estimator runs in every mode — nothing about the algorithm changes. Only the geometry does. Watch the readout turn from healthy to pathological, and watch the second reconstruction switch from "impossible" to "just as good as the first".

Each mode below builds a synthetic two-camera scene with a well-defined coincidence, projects the points, adds a little pixel noise, and then re-runs the full estimator live: a normalized eight-point fit for F, a calibrated promotion to E, and a second fit from the next-smallest null direction of the constraint matrix. The diagnostics are computed from scratch on every mode change.

The rig in its own world frame: both camera centers, the scene points, and the two back-projected rays to a sample correspondence. Pure rotation puts both centers at the same place; a small baseline makes the two rays nearly parallel.

Image 2. Dots are the matched pixels; solid lines are epipolar lines from the fitted F, dashed lines from the second model. Where the dashed lines stop passing through the dots, no second reconstruction exists.

Two reconstructions in the same camera-1 frame, each centered and scaled to fit. Grey = model 1, accent = the second model. When a coincidence lets both fit, the clouds are clearly different yet reproject to the same pixels.

camera 1 camera 2 model 1 (F) model 2 (F₂)
⚠ What "second reconstruction" means here. The estimator's first answer is the smallest-null direction of the constraint matrix. The second model is the next-smallest direction — a fully independent fit to the same pixels. On a general scene that direction does not satisfy the data, and its reprojection residual is enormous. On a planar or critical configuration it does satisfy the data, so a second, different 3D story fits just as well: the empirical signature of the ambiguity.
✓

Cheat sheet

Recap

ConfigurationGeometric coincidenceSymptomDetection method
Planar sceneAll points on one plane; correspondences explained by a homography HF not uniquely determined; epipolar lines ill-determined and noise-sensitive; road to a two-fold ambiguityFit H and F; compare GRIC or RANSAC inlier counts; if H is nearly as good, treat the pair as planar
Pure rotationCamera centers coincide (t = 0); motion is a homographyE = [t]×R collapses; no baseline, so triangulation is impossible; every depth is unobservableCheck the recovered translation norm / baseline; triangulation angle ≈ 0; E's singular spectrum
Small baseline (near-degenerate)Ray-intersection angle shrinks toward zero; rays nearly parallelEnormous, elongated depth uncertainty ΔZ ≈ σ·Z²/(fB) even with tiny reprojection errorMedian triangulation angle vs a threshold; baseline-to-depth ratio; per-point covariance / uncertainty ellipsoid
Critical surfaceAll points on a ruled quadric through both camera centers (hyperboloid, cylinder through the baseline, or plane)Exact two-fold ambiguity: two distinct reconstructions (cameras and points) reproject identically; residuals equalSecond model fits equally well; two candidate reconstructions with near-identical residual; degeneracy of the essential matrix
Repeated / symmetric structureLocally identical texture at many scene locationsAliased feature matches that look like ordinary outliers; a geometrically consistent but wrong model can winDescriptor ratio test; geometric-consistency checks; RANSAC inlier distribution; cycle/loop consistency
Forward motionTranslation mostly along the optical axis; peripheral parallax vanishesNear-degenerate F; small triangulation angle far from the focus of expansion; weak leverage against aliased matchesMedian parallax angle; baseline direction vs optical axis; cheirality and depth-sign checks
Degenerate geometry is where two-view estimation stops being a well-posed algebra problem and becomes a modeling decision. The final part steps back from the classical pipeline entirely — to SLAM, dense stereo, and the learned successors that replace the constraints with predictions. Continue: beyond multi-view geometry →