Degenerate Configurations
Every estimator in this series — the eight-point algorithm, homography DLT, triangulation, PnP — is a constraint solver. It stacks observations into a matrix and trusts that matrix to have a unique, well-conditioned answer. That trust is not a theorem about cameras; it is an assumption about the scene. Flat walls, pure panning, drone footage with a hair-thin baseline, and symmetric buildings all break it. When they do, the solver does not throw an error — it returns a confident-looking wrong answer. This part collects the failure modes in one place, and shows the checks real pipelines use to catch them.
An estimator is only as good as the geometry feeding it
Motivation
Look back at the shape of every estimation step so far. The eight-point algorithm builds a matrix A from correspondences and takes its smallest singular vector. Homography DLT does the same with a different A. Triangulation solves a small linear system; PnP-RANSAC repeatedly solves a minimal one. In each case the answer is the null space of a matrix — and a null space is only a single, stable direction when the geometric configuration that built the matrix is generic.
A degenerate configuration is a specific, nameable coincidence in the scene and camera placement where that stops being true. There are exactly two ways it can go wrong. Either the constraint matrix loses rank — a second direction also nearly satisfies every row, so the solver's choice of "the" answer becomes arbitrary and noise-dominated. Or two genuinely different 3D explanations fit the same images equally well — a two-fold or continuous ambiguity that no amount of residual minimization can resolve, because both explanations have the same residual.
The important practical fact is that these cases do not announce themselves. A degenerate fit still returns an F, still gives a reprojection error near zero on the (degenerate) input, and still looks like a successful reconstruction. What distinguishes a robust pipeline from a naive one is that it checks for the coincidences this page describes, and falls back to a different model when it finds one.
A gallery of degeneracies
Foundations
Each entry below is a configuration that has been studied on its own in this series. Here they are side by side, each stated as a geometric coincidence and the symptom it produces.
Planar scene
Coincidence. Every matched point lies on a single 3D plane. Two views of a plane are related by a homography x₂ ≅ H x₁, so the correspondences are already fully explained by a 3×3 matrix with 8 degrees of freedom — strictly less structure than the fundamental matrix F claims to capture.
Symptom. The eight-point system becomes ill-posed: the constraint matrix has a near-degenerate null space, so the fitted F is not uniquely determined and is extremely sensitive to pixel noise. The epipolar lines it produces are still roughly consistent with the training points, but they drift under even a tiny resampling. This is the classical road to a two-fold ambiguity: a plane admits two distinct decompositions, and both can satisfy the same image data.
Pure rotation (no translation)
Coincidence. The camera rotates about its own optical center, so the two camera centers coincide: t = 0.
Symptom. The essential matrix is built as E = [t]×R, and with t = 0 that product is identically zero — there is no epipolar geometry to recover at all. Equivalently, the two views are related by a pure rotation homography, and because the baseline is zero, triangulation is impossible: there is no parallax, so every depth is unobservable. A pipeline that blindly reports a depth map here is reporting prior, not measurement.
Small baseline (near-degenerate)
Coincidence. The two camera centers are separated by a distance that is tiny compared to the scene depth, so the rays that meet at a scene point are nearly parallel.
Symptom. The two rays still intersect, so triangulation is not impossible — it is merely catastrophically ill-conditioned. With baseline B, focal length f and depth Z, a pixel-sized error σ becomes a depth error ΔZ ≈ σ·Z²/(fB): quadratic in distance, and only linearly (inversely) helped by baseline. The triangulated points come out with enormous, elongated covariance along the viewing direction, even though the reprojection residual still looks small.
Critical surfaces
Coincidence. Every 3D point lies on a ruled quadric that passes through both camera centers — a one-sheeted hyperboloid, a circular cylinder through the baseline, or (as the degenerate case) a plane. The camera centers are on the surface, not merely near it.
Symptom. This is the configuration with a genuine, exact two-fold ambiguity: there are two distinct reconstructions — two different camera placements and two different point clouds — that reproject to bit-for-bit the same pixels. Unlike noise or a small baseline, no residual-minimizing method can separate them, because their residuals are equal. The twin is the "twisted pair" reconstruction, and cheirality does not necessarily rule it out, because both twins can place points in front of both cameras.
Repeated / symmetric structure and forward motion
Coincidence. The scene contains repeated texture — windows, tiles, a colonnade — whose local appearance is nearly identical; or the camera moves mostly along its optical axis, so peripheral points show almost no parallax.
Symptom. Two mechanisms compound each other. Repeated structure makes feature matching aliased: a descriptor can match the wrong instance, producing geometrically inconsistent correspondences that look like ordinary outliers. Forward motion then shrinks the epipolar geometry toward degeneracy, exactly as in the small-baseline case, so the geometric check that would normally reject the aliased match has little leverage. The failure mode is a confidently wrong model built from a majority that happens to agree with itself.
Detection and model selection
Foundations
Because a degenerate fit reports a small residual, residual alone is useless as a detector. Robust pipelines instead ask three separate questions: which model is simpler that fits just as well, how well-determined is the constraint matrix, and does the geometry itself have enough parallax to support the answer.
Model selection: comparing H against F
The cleanest test for a planar scene is to fit both a homography (8 DOF) and a fundamental matrix (7 DOF) to the same correspondences and compare. A penalized score does the judging. The classic one is GRIC, the Geometric Robust Information Criterion: a sum of robust squared residuals plus a per-degree-of-freedom penalty λ·d·n (with n the number of points). Its whole point is that a more complex model must earn its extra parameters by explaining the data clearly better. If H is nearly as good as F, GRIC prefers H — which is exactly the correct call on a plane, and it also catches pure rotation, where the motion is a homography. A related, cheaper heuristic is simply to compare RANSAC inlier counts: when the homography model reaches almost the same inlier count as F, suspect planarity.
Conditioning of the constraint matrix
For a well-posed fit, the smallest singular value of A is clearly separated from the next one: the null space is a single, unambiguous direction. For a rank-deficient configuration the two smallest values crowd together, so the "smallest singular vector" is a poorly determined mixture of two candidates. The ratio σ₂/σ₁ (second-smallest over smallest) therefore reads large and healthy on a general scene and near one as the configuration degenerates. This is the single most useful scalar to watch, and it is reported live in the demo below.
A cheap geometric gate: baseline and triangulation angle
Before trusting any per-point depth, measure the angle between the two back-projected rays that meet at the point (equivalently, the parallax). Healthy two-view geometry has a healthy spread of these angles; pure rotation collapses them to zero, and a small baseline collapses them to a fraction of a degree. A median-angle threshold is a fast, model-free way to discard a bad pair or mark a region as depth-unreliable — and it is exactly the quantity the depth-uncertainty formula ΔZ ≈ σ·Z²/(fB) is punishing.
Degeneracy checks on the essential matrix
Once the cameras are calibrated, F is promoted to E = Kᵀ F K, and a valid essential matrix has a very specific spectrum: two equal nonzero singular values and a third that is zero, σ₁ = σ₂ > σ₃ = 0. The shape of that spectrum is itself a diagnostic. If the two nonzero values come out nearly equal, the matrix is genuinely an essential matrix; if one of them collapses toward zero (so the matrix is closer to rank 1), the estimated geometry is sitting in or near a degenerate configuration and should not be trusted as calibrated motion.
Play: switch on a degeneracy and watch the estimator break
Interactive
Each mode below builds a synthetic two-camera scene with a well-defined coincidence, projects the points, adds a little pixel noise, and then re-runs the full estimator live: a normalized eight-point fit for F, a calibrated promotion to E, and a second fit from the next-smallest null direction of the constraint matrix. The diagnostics are computed from scratch on every mode change.
The rig in its own world frame: both camera centers, the scene points, and the two back-projected rays to a sample correspondence. Pure rotation puts both centers at the same place; a small baseline makes the two rays nearly parallel.
Image 2. Dots are the matched pixels; solid lines are epipolar lines from the fitted F, dashed lines from the second model. Where the dashed lines stop passing through the dots, no second reconstruction exists.
Two reconstructions in the same camera-1 frame, each centered and scaled to fit. Grey = model 1, accent = the second model. When a coincidence lets both fit, the clouds are clearly different yet reproject to the same pixels.
Cheat sheet
Recap
| Configuration | Geometric coincidence | Symptom | Detection method |
|---|---|---|---|
| Planar scene | All points on one plane; correspondences explained by a homography H | F not uniquely determined; epipolar lines ill-determined and noise-sensitive; road to a two-fold ambiguity | Fit H and F; compare GRIC or RANSAC inlier counts; if H is nearly as good, treat the pair as planar |
| Pure rotation | Camera centers coincide (t = 0); motion is a homography | E = [t]×R collapses; no baseline, so triangulation is impossible; every depth is unobservable | Check the recovered translation norm / baseline; triangulation angle ≈ 0; E's singular spectrum |
| Small baseline (near-degenerate) | Ray-intersection angle shrinks toward zero; rays nearly parallel | Enormous, elongated depth uncertainty ΔZ ≈ σ·Z²/(fB) even with tiny reprojection error | Median triangulation angle vs a threshold; baseline-to-depth ratio; per-point covariance / uncertainty ellipsoid |
| Critical surface | All points on a ruled quadric through both camera centers (hyperboloid, cylinder through the baseline, or plane) | Exact two-fold ambiguity: two distinct reconstructions (cameras and points) reproject identically; residuals equal | Second model fits equally well; two candidate reconstructions with near-identical residual; degeneracy of the essential matrix |
| Repeated / symmetric structure | Locally identical texture at many scene locations | Aliased feature matches that look like ordinary outliers; a geometrically consistent but wrong model can win | Descriptor ratio test; geometric-consistency checks; RANSAC inlier distribution; cycle/loop consistency |
| Forward motion | Translation mostly along the optical axis; peripheral parallax vanishes | Near-degenerate F; small triangulation angle far from the focus of expansion; weak leverage against aliased matches | Median parallax angle; baseline direction vs optical axis; cheirality and depth-sign checks |