Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

0

Why an uncalibrated reconstruction is only projective

Motivation

Every camera matrix P and world point X only ever meet through the product PX. Squeeze any invertible 4×4 matrix H in between and the image is bit-for-bit unchanged:

x ∝ P X = (P H⁻¹)(H X)

So from correspondences alone — with K unknown for every camera — the best you can ever get is a projective reconstruction: the true scene up to some completely general 4×4 H, not merely a rotation, translation and scale. Angles are wrong, parallel lines need not stay parallel, and ratios of lengths along a line are meaningless. It is a real reconstruction, with real camera poses and real 3D points, but it has never been told where "at infinity" or "perpendicular" live.

Self-calibration claws that back without a calibration object, using one physical fact: the same camera — the same lens, the same sensor, the same K — took all the photos. Each image's own ω must therefore be the same ω, no matter how the camera was moved or pointed. That single consistency condition, applied across many views, is enough to pin down the metric structure of the scene from the images alone.

💡 Two layers of ambiguity, two layers of repair. A projective reconstruction becomes affine once you locate the plane at infinity π∞ (the projective → affine step), and metric once you locate the absolute dual quadric Ω*∞ (the affine → metric step). This page is the second step. The first — locating π∞ — is what the vanishing points of parallel lines, or the line at infinity of a reference plane, are for.
1

The absolute dual quadric Ω*∞

The one object that carries the metric

The absolute conic Ω∞ (Part 1) is the imaginary conic living inside the plane at infinity whose equation is X² + Y² + Z² = 0 with W = 0. Going one dimension up means taking its dual: the set of planes tangent to it, not the points upon it. That dual object is the absolute dual quadric Ω*∞, and algebraically it is simply a symmetric 4×4 matrix.

Its bookkeeping is exactly that of any degenerate quadric. A symmetric 4×4 matrix has 4·5/2 = 10 independent entries; throwing away the overall scale leaves 10 degrees of freedom, and because it is the dual of a conic rather than a genuine 3D surface it is rank 3 — degenerate, with a single zero eigenvalue. Like every object in projective geometry it is defined only up to scale, Ω*∞ ∝ λΩ*∞.

Written in the true, metric frame, Ω*∞ has a dazzlingly simple form:

Ω*∞ = diag(1, 1, 1, 0)    (in the metric frame)

That bare matrix encodes every right angle and every ratio of lengths in the world, and it is exactly the object a projective reconstruction has lost track of. The reason it is the right target is the imaging relation. Put P = K[R | t] and push Ω*∞ through the camera:

P Ω*∞ Pᵗ = K[R | t] diag(1,1,1,0) [R | t]ᵗ Kᵗ
      = K R Rᵗ Kᵗ = K Kᵗ

The translation column is multiplied by zero and the rotation cancels against its own transpose, leaving K Kᵗ — the dual image of the absolute conic ω*. This is the central equation of self-calibration:

ω* = K Kᵗ ∝ P Ω*∞ Pᵗ    and equivalently    ω = (K Kᵗ)⁻¹ ∝ (P Ω*∞ Pᵗ)⁻¹ = K⁻ᵗK⁻¹

It says that a camera's own ω — the very conic the calibration page estimates — is nothing but the shadow of one shared 4×4 object. Unknown K per camera has become one unknown Ω*∞ for the whole set. That is the lever: instead of a different K per image, we solve for a single global quadric whose image in every camera must simultaneously reproduce that camera's ω.

What makes the lever practical is that P Ω*∞ Pᵗ is linear in the ten entries of Ω*∞. Once we commit to a few simple facts about K, each camera hands us a small set of linear equations in those ten numbers, and the whole problem collapses from a nonlinear search to a least-squares solve. Step 2 lists exactly which assumptions buy which equations.

💡 Consistency with Parts 1 and 5: there, ω = K⁻ᵗK⁻¹ with dual ω* = KKᵗ. Nothing changes here. Self-calibration just adds a global object Ω*∞ whose image under each camera is that same ω*, so every part of the series is solving for the same conic by a different route.
2

Kruppa's equations, and the modern linear replacement

From nonlinear pair constraints to a linear solve

The classical route, due to Kruppa, looks at two cameras at a time. Two views are tied together by their epipolar geometry F (Part 6), and each view carries the same image conic ω. But F and ω cannot be chosen independently: F already encodes the relative pose, and the scene's right angles are buried in ω, so the two must be mutually consistent. Forcing that consistency yields the Kruppa equations: a pair of quadratic constraints linking the entries of F to those of ω, one pair per camera pair.

Those equations are correct and historically decisive, but they are awkward in practice. They are quadratic, they only ever involve a pair at a time, and with noisy correspondences they are badly conditioned — a small error in F can swing the recovered ω violently. Modern self-calibration therefore replaces them with a global, linear formulation built on Ω*∞:

1. Stack the constraints  P_i Ω*∞ P_iᵗ ∝ ω* = K Kᵗ  across every view into one linear system in the ten entries of Ω*∞
2. Solve by least squares (the null space of the stacked system)
3. Project the answer back onto the rank-3 cone — Ω*∞ must be degenerate, and this step is what removes the last spurious direction

Because every stage is linear algebra except the final rank-3 projection, the whole thing is far better conditioned than the Kruppa equations, and it uses all views at once rather than one pair at a time.

What each assumption buys. A general symmetric ω* has 5 degrees of freedom, but the proportionality P_i Ω*∞ P_iᵗ ∝ K Kᵗ alone is not enough to pin down ten unknowns. Every simplifying assumption about K removes unknowns and turns the relation into more linear equations:

• known principal point (c_x, c_y)  →  the bottom row of ω* is known
• zero skew  →  ω*[0,1] = c_x c_y is known
• unit aspect ratio f_x = f_y  →  ω*[0,0] − ω*[1,1] = c_x² − c_y² is known
• the same K across all views  →  all N_i must agree on the one remaining unknown, f

Under the three calibration assumptions above — known principal point, zero skew, unit aspect — every view contributes four equations that are linear in the entries of Ω*∞ and involve no focal length at all. Writing N_i = P_i Ω*∞ P_iᵗ:

N_i[0,1] − c_x c_y N_i[2,2] = 0
N_i[0,2] − c_x N_i[2,2] = 0
N_i[1,2] − c_y N_i[2,2] = 0
N_i[0,0] − N_i[1,1] − (c_x² − c_y²) N_i[2,2] = 0

Ten unknowns, up to an overall scale, is nine degrees of freedom, and four equations per view means three or more views overdetermine the system. Two views leave it short; the more views, the more the extra rows average out numerical error and the better conditioned the solve becomes. The fewer assumptions you are willing to make, the more of these equations disappear and the more views you need to make up the difference — the trade at the heart of every self-calibration method.

Once Ω*∞ is known, K follows immediately. For any view, N_i ∝ K Kᵗ, so after normalizing N_i by its [2,2] entry the focal length is read straight off the diagonal:

f² = N_i[0,0]/N_i[2,2] − c_x² = N_i[1,1]/N_i[2,2] − c_y²    (equal by the unit-aspect assumption)
3

Play: solve for ω and watch the metric upgrade converge

Interactive

🎯 Learning goal: feed the solver nothing but projectively-distorted camera matrices and the three calibration assumptions, and watch it claw back the true focal length — and with it the right angles of the scene. Add views and the recovered f tightens onto the truth; at two views the linear system is simply not yet determined.

The setup is entirely synthetic and entirely honest. A ground-truth camera has K = diag(f, f, 1) with the principal point fixed at a known image center, zero skew and unit aspect ratio; its focal length f is the one thing we will pretend not to know. Several cameras sit at genuinely three-dimensional positions and photograph a cloud of 3D points. We then apply a random projective transformation H_proj to the whole scene at once — every camera becomes P_i H_proj⁻¹, every point becomes H_proj X — which leaves every image pixel exactly unchanged (the demo measures and prints that reprojection error) while destroying all metric information. That distorted camera set is the "uncalibrated reconstruction."

The solver then does exactly what Step 2 describes. It assumes a shared K with the known principal point, zero skew and unit aspect, builds the four linear constraints P_i Ω*∞ P_iᵗ ∝ K Kᵗ for every view, solves the homogeneous least-squares system for the ten entries of Ω*∞, and imposes the rank-3 condition. Because the data are exact, the recovered f should match the ground truth to many digits; because the linear system is genuinely underdetermined below three views, the two-view setting is reported as unsolved rather than faked.

Left: a 2 × 1 rectangle corner from the scene, as it looks in the projective reconstruction (the angle at the corner is not 90°). Right: the same corner after the metric upgrade recovered from Ω*∞ — the right angle and the 2:1 edge ratio are restored. Each panel is drawn in its own plane's best-fit basis, so the picture shows the true in-plane shape.

projective (H_proj applied) after metric upgrade
⚠ What's simplified here, plainly: the correspondences and the distorted reconstruction are exact (no pixel noise, no RANSAC, no real feature matching), so the constraints could in principle be solved to machine precision; the demo replaces a real projective reconstruction with the direct application of a random H_proj, which gives the identical kind of distorted cameras without re-deriving them from matches. The rank-3 condition is imposed by searching the two-dimensional null space of the linear system for the solution on which all views agree on a common focal length, then zeroing the smallest eigenvalue — a self-contained version of the standard "linear solve, then project to rank 3" recipe. With exact data the recovered f lands on the true value to many digits from three views upward, and the error keeps shrinking as views are added; below three views the system is underdetermined and is reported as such. Nothing here is scripted or back-solved.
✓

Cheat sheet

Recap

PieceWhat it is
Absolute dual quadric Ω*∞Symmetric 4×4 matrix, defined up to scale, 10 DOF, rank 3 (degenerate). The dual of the absolute conic Ω∞ one dimension up; in the metric frame Ω*∞ = diag(1,1,1,0).
Imaging relationP Ω*∞ Pᵗ ∝ ω* = K Kᵗ, equivalently ω = K⁻ᵗK⁻¹ ∝ (P Ω*∞ Pᵗ)⁻¹. The camera's own ω is the shadow of one shared quadric.
Image of the absolute conicω = K⁻ᵗK⁻¹, dual ω* = KKᵗ; 5 DOF, depending on K alone — the same conic as Parts 1 and 5.
Kruppa equationsClassical pair-wise constraints: for each camera pair, F and ω must be mutually consistent, giving quadratic equations in ω. Correct but poorly conditioned.
Modern linear formReplace Kruppa with linear constraints on Ω*∞ stacked across all views, solve by least squares, then project to rank 3. Better conditioned and uses every view at once.
Linear constraints (assumed K)With known principal point, zero skew, unit aspect, each view adds four equations linear in Ω*∞: N[0,1]−c_x c_y N[2,2]=0, N[0,2]−c_x N[2,2]=0, N[1,2]−c_y N[2,2]=0, N[0,0]−N[1,1]−(c_x²−c_y²)N[2,2]=0.
View requirementsFour equations × V views for nine degrees of freedom. Two views are not enough; three views overdetermine it, and adding views improves conditioning and accuracy. Fewer assumptions → more views needed.
Recovering KFrom any view, N_i = P_i Ω*∞ P_iᵗ ∝ KKᵗ; normalize by N_i[2,2], then f² = N_i[0,0]/N_i[2,2] − c_x² = N_i[1,1]/N_i[2,2] − c_y².
Relation to Zhang's calibrationZhang estimates ω from images of a known planar target; self-calibration estimates Ω*∞ (hence ω, hence K) from unknown scene and motion. Both end by factoring the image conic to recover K.
Relation to the stratified upgradeSupplies the affine → metric step. The projective → affine step — locating the plane at infinity π∞ — comes first; Ω*∞ then fixes angles and length ratios.
Self-calibration recovers the metric structure whenever the camera assumptions hold and the motion is general enough. Some camera motions — pure rotation, or a camera sliding along a line — leave the metric structure genuinely unobservable, which is where the next part goes. Continue: degenerate configurations →