Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

0

The problem, and the object that's actually linear

Setup

Calibration means recovering the intrinsic matrix K — and often lens distortion coefficients — from images of a target whose geometry you know exactly, almost always a flat checkerboard, photographed from several different poses (different positions and tilts relative to the camera). The obstacle is that K doesn't sit anywhere you can solve for linearly: it shows up inside a perspective divide, buried in a product with an unknown rotation and translation for every single photo. Trying to fit fx, fy, cx, cy directly against pixel measurements is a nasty nonlinear problem with a different unknown pose per image.

The fix is to stop looking for K and look for a related quantity that is linear: the image of the absolute conic, written ω (omega):

ω = K⁻ᵀK⁻¹

Part 1 introduced a conic as a symmetric 3×3 matrix C with x̃ᵀCx̃ = 0. ω is exactly that kind of object — the image, under this camera, of a specific conic that lives entirely at infinity in 3D (the "absolute conic") and is never itself visible in any photo. What makes it useful isn't the geometric picture though, it's two algebraic facts: ω is symmetric (6 numbers, and since it's only ever used up to scale, 5 real degrees of freedom — the same 5 DOF as a general K), and it depends only on K, never on where the camera is pointed. A camera held at ten different angles has ten different extrinsics but the exact same ω in every shot. That's the whole strategy: each photo of the checkerboard contributes linear equations in the six entries of ω, and once enough photos have piled up enough equations, solving for ω is ordinary linear algebra.

Getting from ω back to K is also linear algebra, via a Cholesky decomposition. Factor the symmetric positive-definite matrix ω = LLᵀ with L lower-triangular. Comparing to the definition ω = K⁻ᵀK⁻¹ = (K⁻¹)ᵀ(K⁻¹) and matching factors gives K⁻¹ = Lᵀ — so K is just the inverse of the upper-triangular matrix Lᵀ, rescaled so its bottom-right entry is 1 (the usual homogeneous normalization).

ω = LLᵀ  (Cholesky)  ⟹  K⁻¹ = Lᵀ  ⟹  K = (Lᵀ)⁻¹,  normalized so K₃₃ = 1
💡 The plan for the rest of this page: get linear constraints on ω from checkerboard homographies (Zhang's method), stack constraints from several poses and solve for ω with the same homogeneous-least-squares-by-SVD trick used elsewhere in this series, Cholesky down to K, then refine everything nonlinearly against actual pixel error.
1

Zhang's method: from a flat checkerboard to constraints on ω

Derivation

Put the checkerboard's own coordinate frame on the board itself, so every corner has Z = 0. A board-plane point (X, Y, 0) projects to a pixel by the usual x̃ = K[r₁ r₂ r₃ t](X, Y, 0, 1)ᵀ — but the Z = 0 kills the r₃ column entirely, leaving a plain 3×3 homography:

x̃ ≅ H (X, Y, 1)ᵀ,    H ≅ K[r₁  r₂  t]

That H is estimable directly from board-corner ↔ pixel correspondences in a single image (that's ordinary homography fitting, the subject of Part 8) — one homography per photo, no camera geometry needed yet. Write its three columns as h₁, h₂, h₃. Since H ≅ K[r₁ r₂ t] only up to an unknown overall scale λ (the usual homogeneous ambiguity — the K in front doesn't fix the scale of H's columns), matching columns gives:

h₁ = λ K r₁  ⟹  r₁ = (1/λ) K⁻¹h₁      h₂ = λ K r₂  ⟹  r₂ = (1/λ) K⁻¹h₂

r₁ and r₂ are two columns of a rotation matrix, so they're orthonormal: r₁ᵀr₂ = 0 and r₁ᵀr₁ = r₂ᵀr₂. Substitute the expressions above — the unknown λ cancels out of both, and K⁻ᵀK⁻¹ is exactly ω:

r₁ᵀr₂ = (1/λ²) h₁ᵀ K⁻ᵀK⁻¹ h₂ = (1/λ²) h₁ᵀωh₂ = 0    ⟹    h₁ᵀωh₂ = 0
r₁ᵀr₁ = r₂ᵀr₂  ⟹  (1/λ²)h₁ᵀωh₁ = (1/λ²)h₂ᵀωh₂    ⟹    h₁ᵀωh₁ = h₂ᵀωh₂

Both are linear in the six unknown entries of ω — a quadratic form hᵀωh' expands into a fixed linear combination of {ω₁₁, ω₁₂, ω₂₂, ω₁₃, ω₂₃, ω₃₃} with coefficients built only from the known entries of h and h'. Stack ω into a 6-vector b and each constraint becomes vᵀb = 0 for a known row-vector v:

b = (ω₁₁, ω₁₂, ω₂₂, ω₁₃, ω₂₃, ω₃₃)ᵀ    v₂ = h₁ᵀωh₂ = 0    (v₁₁ − v₂₂)ᵀb = h₁ᵀωh₁ − h₂ᵀωh₂ = 0

Every checkerboard pose contributes exactly 2 rows. b has 5 true degrees of freedom (6 numbers, minus 1 for the overall scale ambiguity that homogeneous quantities always carry), so 2 equations per pose means you need at least 5 / 2 = 2.5 poses — rounded up, 3 poses in generically different orientations, to pin b down. Stack the rows from all n poses into one 2n × 6 matrix V and solve the homogeneous system Vb = 0 — the exact same trick as the eight-point algorithm elsewhere in this series: the least-squares solution (in the presence of noise, V won't be exactly singular) is the right singular vector of V with the smallest singular value, equivalently the eigenvector of VᵀV with the smallest eigenvalue. The demo below finds it with a small Jacobi eigenvalue solver written from scratch, on the 6×6 matrix VᵀV.

From there: reshape b into ω, Cholesky-decompose to get K̂ (previous section), then peel off each pose's extrinsics too — λ = 1/‖K̂⁻¹h₁‖, r₁ = λK̂⁻¹h₁, r₂ = λK̂⁻¹h₂, r₃ = r₁×r₂, t = λK̂⁻¹h₃. That's a full closed-form calibration from a handful of photos — but it's still only a linear approximation: it minimizes an algebraic residual on V, not actual pixel reprojection error, and it says nothing about lens distortion. Real pipelines always follow it with nonlinear refinement — Levenberg–Marquardt, jointly optimizing K, distortion coefficients, and every single pose's R and t together, minimizing true reprojection error in pixels. Closed-form linear init, then nonlinear polish — the same two-stage pattern you'll see again for triangulation, PnP, and bundle adjustment later in this series.

⚠️ What the demo below skips: estimating H itself from noisy pixel correspondences (that's homography fitting via DLT, Part 8's topic) and the nonlinear refinement step. It starts from exact per-pose homographies and solves only the ω-recovery step above — enough to see why pose diversity, not pose count, is what actually makes the linear system solvable.
2

Play: the calibration lab

Interactive

🎯 Learning goal: add checkerboard poses one at a time and watch the recovered K̂ converge toward the true K — then try the degenerate preset and watch the same math fail in exactly the way real calibration sessions fail.

A hidden camera with a known, fixed K is watching a checkerboard. Each button below adds one more pose of that board (a homography H = K[r₁ r₂ t] built from a real rotation and translation) to the pool. Every time the pool changes, the page builds V from all current poses, solves Vb = 0 for ω by Jacobi eigen-decomposition, Cholesky-decomposes to K̂, and reprojects the board corners with the recovered K̂ and per-pose extrinsics to report an RMS pixel error.

The pose pool in 3D (camera frame): the camera is the small pyramid at the origin, each colored grid is one checkerboard pose. Fronto-parallel poses all share one board normal — the geometric reason they’re degenerate.

What the calibration software sees: detected corners of each pose (dots + grid), with pixel noise and lens distortion applied. H is fitted from these observations by DLT.

Add Tilt A, Tilt B, Tilt C (three genuinely different orientations) and watch K̂ lock onto the true K. With zero noise the reprojection error is numerically ~0; drag corner noise up to 1–2 px and it lands in the realistic 0.1–1 px band that Step 5's thresholds are about. Now hit Reset and instead add several Fronto-parallel poses: the board always faces the camera square-on, only sliding around and spinning in-plane between shots. Every fronto-parallel pose has the same plane normal in the camera's frame (see the 3D view), so their v rows keep pointing in nearly the same directions no matter how many you stack — V stays rank-deficient, and the recovered ω is frequently not even positive-definite, so the Cholesky step itself fails outright.

Try the barrel distortion slider with a few well-posed tilts: a single radial term breaks the projective assumption behind Zhang's linear solve — watch the pinhole-only fit RMS climb well past 1 px no matter how many poses you add, with K̂ visibly biased. The lab then grid-searches k̂₁ (the same reprojection-error objective a joint optimizer refines), undistorts the corners, and re-runs the linear solve — RMS drops back to the noise floor and K̂ snaps back toward truth. That pinhole-fails-then-radial-refit sequence is exactly how real calibration pipelines work, one extra parameter at a time.

3

The other two degrees of freedom: skew and non-square pixels

Extending K

Part 4's demo used the simplest possible K: a single shared focal length (fx = fy) and zero skew. The fully general intrinsic matrix has 5 degrees of freedom, and the Zhang derivation above never assumed anything simpler:

K = fxscx 0fycy 001

Skew s is nonzero when the sensor's row and column axes aren't exactly perpendicular. On virtually every real manufactured sensor this is extremely close to zero — it's rarely worth modeling — but the calibration math has to allow for it to be correct in general, and Zhang's linear solve recovers it automatically as one of ω's five degrees of freedom with no extra work. fx ≠ fy means pixels aren't square: either the physical sensor elements aren't square (common on older or cheaper sensors), or — far more common in practice — the image was resized or cropped anisotropically somewhere in the pipeline (different scale factors in x and y), which silently turns square pixels into rectangular ones from the math's point of view.

⚠️ Skew is small enough on real hardware that plenty of practical calibration pipelines fix s = 0 by assumption to make the nonlinear refinement step better-conditioned. That's a modeling choice, not a mathematical necessity — the demo below keeps the full 5-DOF form so you can see what skew actually does to an image.
4

Play: the full 5-DOF K

Interactive

🎯 Learning goal: skew shears the image, fx ≠ fy stretches it unevenly — both are ordinary entries of K, computed live from x̃ = K[R|t]X̃ with a fixed camera pose.

A checkerboard at a fixed, tilted pose, projected through a fully general K. The two arrows show where K sends the camera-plane's unit x- and y-directions — they're literally K's first two columns. When they're perpendicular and equal length, pixels are square and skew is zero, same as every earlier demo in this series.

Push skew away from 0 and the checkerboard squares turn into parallelograms — the grid lines that were vertical in the world stop being vertical in the image, even though nothing about the camera's pose moved. Separate fx and fy and the squares stretch into rectangles along one axis. Both effects are exactly what the two extra off-diagonal-ish degrees of freedom in K encode.

5

What a good calibration looks like, and how it fails

Diagnosis

After nonlinear refinement, the number to look at is RMS reprojection error: reproject every checkerboard corner from every pose using the fitted K, distortion, and per-pose extrinsics, and measure the distance in pixels to where the corner detector actually found it. For a decent lens and a reasonably sharp checkerboard, a good calibration lands well under 0.5 px. Anything creeping above 1 px means something is wrong — a mis-measured square size, a checkerboard that flexed or moved mid-capture, a distortion model that's the wrong shape for the lens, or (the case this page's demo makes concrete) checkerboard poses that don't actually constrain the problem.

That last failure mode is common and easy to walk into by accident: if every pose is coplanar with the others in a way that doesn't vary enough — the board is always held roughly fronto-parallel to the camera, or every pose is a rotation about the same axis — the linear system for ω becomes rank-deficient, exactly as the degenerate preset in the demo above shows. The fix that every calibration guide repeats ("tilt the board, cover the whole frame, rotate it around different axes between shots") isn't superstition — it's making sure the v rows stacked into V actually span enough directions for Vb = 0 to pin down a unique b. More poses without more orientation diversity doesn't help; three well-chosen tilts beat thirty photos of the board held flat.

✓

Cheat sheet

Recap

ObjectSymbol / formDefinitionWhere it matters
Image of the absolute conicω = K⁻ᵀK⁻¹Symmetric 3×3, 5 DOF up to scale; same for every pose of one cameraThe thing Zhang's method actually solves for
K from ωω = LLᵀ, K = (Lᵀ)⁻¹Cholesky decomposition, normalized so K₃₃ = 1Closed-form step after the linear solve
Zhang constraint 1h₁ᵀωh₂ = 0From r₁ᵀr₂ = 0 (rotation columns orthogonal)One row of V per pose
Zhang constraint 2h₁ᵀωh₁ = h₂ᵀωh₂From r₁ᵀr₁ = r₂ᵀr₂ (rotation columns unit length)Second row of V per pose
Minimum poses≥ 3, orientation-diverse2 constraints/pose, 5 unknowns up to scale ⟹ ≥2.5 posesFewer, or all-coplanar, ⟹ V rank-deficient
Full intrinsicsK = [[fx,s,cx],[0,fy,cy],[0,0,1]]5 DOF: focal lengths, skew, principal pointPart 4's demo fixed fx=fy, s=0
Reprojection errorRMS, in pixelsDistance between reprojected and detected corners< 0.5 px good, > 1 px investigate
With K in hand, it's time to bring in a second camera. Two cameras looking at the same point are constrained to each other in a very specific way — that constraint is epipolar geometry. Continue: epipolar geometry →