Camera Calibration
Every part of this series so far has quietly said "assume the camera is calibrated" and jumped straight to a tidy K = diag(f, f) with the principal point near the image center. Nobody hands you that matrix. In practice you print a checkerboard of known size, wave it in front of the lens a handful of times, and estimate K — and usually lens distortion too — from those photos alone. This page derives how that estimation actually works: a quantity that's linear in the unknowns even though K itself isn't, a closed-form solve from a few checkerboard poses (Zhang's method), and the same "cheap linear init, then polish with nonlinear optimization" pattern that shows up everywhere else in multi-view geometry.
The problem, and the object that's actually linear
Setup
Calibration means recovering the intrinsic matrix K — and often lens distortion coefficients — from images of a target whose geometry you know exactly, almost always a flat checkerboard, photographed from several different poses (different positions and tilts relative to the camera). The obstacle is that K doesn't sit anywhere you can solve for linearly: it shows up inside a perspective divide, buried in a product with an unknown rotation and translation for every single photo. Trying to fit fx, fy, cx, cy directly against pixel measurements is a nasty nonlinear problem with a different unknown pose per image.
The fix is to stop looking for K and look for a related quantity that is linear: the image of the absolute conic, written ω (omega):
Part 1 introduced a conic as a symmetric 3×3 matrix C with x̃ᵀCx̃ = 0. ω is exactly that kind of object — the image, under this camera, of a specific conic that lives entirely at infinity in 3D (the "absolute conic") and is never itself visible in any photo. What makes it useful isn't the geometric picture though, it's two algebraic facts: ω is symmetric (6 numbers, and since it's only ever used up to scale, 5 real degrees of freedom — the same 5 DOF as a general K), and it depends only on K, never on where the camera is pointed. A camera held at ten different angles has ten different extrinsics but the exact same ω in every shot. That's the whole strategy: each photo of the checkerboard contributes linear equations in the six entries of ω, and once enough photos have piled up enough equations, solving for ω is ordinary linear algebra.
Getting from ω back to K is also linear algebra, via a Cholesky decomposition. Factor the symmetric positive-definite matrix ω = LLᵀ with L lower-triangular. Comparing to the definition ω = K⁻ᵀK⁻¹ = (K⁻¹)ᵀ(K⁻¹) and matching factors gives K⁻¹ = Lᵀ — so K is just the inverse of the upper-triangular matrix Lᵀ, rescaled so its bottom-right entry is 1 (the usual homogeneous normalization).
Zhang's method: from a flat checkerboard to constraints on ω
Derivation
Put the checkerboard's own coordinate frame on the board itself, so every corner has Z = 0. A board-plane point (X, Y, 0) projects to a pixel by the usual x̃ = K[r₁ r₂ r₃ t](X, Y, 0, 1)ᵀ — but the Z = 0 kills the r₃ column entirely, leaving a plain 3×3 homography:
That H is estimable directly from board-corner ↔ pixel correspondences in a single image (that's ordinary homography fitting, the subject of Part 8) — one homography per photo, no camera geometry needed yet. Write its three columns as h₁, h₂, h₃. Since H ≅ K[r₁ r₂ t] only up to an unknown overall scale λ (the usual homogeneous ambiguity — the K in front doesn't fix the scale of H's columns), matching columns gives:
r₁ and r₂ are two columns of a rotation matrix, so they're orthonormal: r₁ᵀr₂ = 0 and r₁ᵀr₁ = r₂ᵀr₂. Substitute the expressions above — the unknown λ cancels out of both, and K⁻ᵀK⁻¹ is exactly ω:
r₁ᵀr₁ = r₂ᵀr₂ ⟹ (1/λ²)h₁ᵀωh₁ = (1/λ²)h₂ᵀωh₂ ⟹ h₁ᵀωh₁ = h₂ᵀωh₂
Both are linear in the six unknown entries of ω — a quadratic form hᵀωh' expands into a fixed linear combination of {ω₁₁, ω₁₂, ω₂₂, ω₁₃, ω₂₃, ω₃₃} with coefficients built only from the known entries of h and h'. Stack ω into a 6-vector b and each constraint becomes vᵀb = 0 for a known row-vector v:
Every checkerboard pose contributes exactly 2 rows. b has 5 true degrees of freedom (6 numbers, minus 1 for the overall scale ambiguity that homogeneous quantities always carry), so 2 equations per pose means you need at least 5 / 2 = 2.5 poses — rounded up, 3 poses in generically different orientations, to pin b down. Stack the rows from all n poses into one 2n × 6 matrix V and solve the homogeneous system Vb = 0 — the exact same trick as the eight-point algorithm elsewhere in this series: the least-squares solution (in the presence of noise, V won't be exactly singular) is the right singular vector of V with the smallest singular value, equivalently the eigenvector of VᵀV with the smallest eigenvalue. The demo below finds it with a small Jacobi eigenvalue solver written from scratch, on the 6×6 matrix VᵀV.
From there: reshape b into ω, Cholesky-decompose to get K̂ (previous section), then peel off each pose's extrinsics too — λ = 1/‖K̂⁻¹h₁‖, r₁ = λK̂⁻¹h₁, r₂ = λK̂⁻¹h₂, r₃ = r₁×r₂, t = λK̂⁻¹h₃. That's a full closed-form calibration from a handful of photos — but it's still only a linear approximation: it minimizes an algebraic residual on V, not actual pixel reprojection error, and it says nothing about lens distortion. Real pipelines always follow it with nonlinear refinement — Levenberg–Marquardt, jointly optimizing K, distortion coefficients, and every single pose's R and t together, minimizing true reprojection error in pixels. Closed-form linear init, then nonlinear polish — the same two-stage pattern you'll see again for triangulation, PnP, and bundle adjustment later in this series.
Play: the calibration lab
Interactive
A hidden camera with a known, fixed K is watching a checkerboard. Each button below adds one more pose of that board (a homography H = K[r₁ r₂ t] built from a real rotation and translation) to the pool. Every time the pool changes, the page builds V from all current poses, solves Vb = 0 for ω by Jacobi eigen-decomposition, Cholesky-decomposes to K̂, and reprojects the board corners with the recovered K̂ and per-pose extrinsics to report an RMS pixel error.
The pose pool in 3D (camera frame): the camera is the small pyramid at the origin, each colored grid is one checkerboard pose. Fronto-parallel poses all share one board normal — the geometric reason they’re degenerate.
What the calibration software sees: detected corners of each pose (dots + grid), with pixel noise and lens distortion applied. H is fitted from these observations by DLT.
Add Tilt A, Tilt B, Tilt C (three genuinely different orientations) and watch K̂ lock onto the true K. With zero noise the reprojection error is numerically ~0; drag corner noise up to 1–2 px and it lands in the realistic 0.1–1 px band that Step 5's thresholds are about. Now hit Reset and instead add several Fronto-parallel poses: the board always faces the camera square-on, only sliding around and spinning in-plane between shots. Every fronto-parallel pose has the same plane normal in the camera's frame (see the 3D view), so their v rows keep pointing in nearly the same directions no matter how many you stack — V stays rank-deficient, and the recovered ω is frequently not even positive-definite, so the Cholesky step itself fails outright.
Try the barrel distortion slider with a few well-posed tilts: a single radial term breaks the projective assumption behind Zhang's linear solve — watch the pinhole-only fit RMS climb well past 1 px no matter how many poses you add, with K̂ visibly biased. The lab then grid-searches k̂₁ (the same reprojection-error objective a joint optimizer refines), undistorts the corners, and re-runs the linear solve — RMS drops back to the noise floor and K̂ snaps back toward truth. That pinhole-fails-then-radial-refit sequence is exactly how real calibration pipelines work, one extra parameter at a time.
The other two degrees of freedom: skew and non-square pixels
Extending K
Part 4's demo used the simplest possible K: a single shared focal length (fx = fy) and zero skew. The fully general intrinsic matrix has 5 degrees of freedom, and the Zhang derivation above never assumed anything simpler:
Skew s is nonzero when the sensor's row and column axes aren't exactly perpendicular. On virtually every real manufactured sensor this is extremely close to zero — it's rarely worth modeling — but the calibration math has to allow for it to be correct in general, and Zhang's linear solve recovers it automatically as one of ω's five degrees of freedom with no extra work. fx ≠ fy means pixels aren't square: either the physical sensor elements aren't square (common on older or cheaper sensors), or — far more common in practice — the image was resized or cropped anisotropically somewhere in the pipeline (different scale factors in x and y), which silently turns square pixels into rectangular ones from the math's point of view.
s = 0 by assumption to make the nonlinear refinement step better-conditioned. That's a modeling choice, not a mathematical necessity — the demo below keeps the full 5-DOF form so you can see what skew actually does to an image.Play: the full 5-DOF K
Interactive
A checkerboard at a fixed, tilted pose, projected through a fully general K. The two arrows show where K sends the camera-plane's unit x- and y-directions — they're literally K's first two columns. When they're perpendicular and equal length, pixels are square and skew is zero, same as every earlier demo in this series.
Push skew away from 0 and the checkerboard squares turn into parallelograms — the grid lines that were vertical in the world stop being vertical in the image, even though nothing about the camera's pose moved. Separate fx and fy and the squares stretch into rectangles along one axis. Both effects are exactly what the two extra off-diagonal-ish degrees of freedom in K encode.
What a good calibration looks like, and how it fails
Diagnosis
After nonlinear refinement, the number to look at is RMS reprojection error: reproject every checkerboard corner from every pose using the fitted K, distortion, and per-pose extrinsics, and measure the distance in pixels to where the corner detector actually found it. For a decent lens and a reasonably sharp checkerboard, a good calibration lands well under 0.5 px. Anything creeping above 1 px means something is wrong — a mis-measured square size, a checkerboard that flexed or moved mid-capture, a distortion model that's the wrong shape for the lens, or (the case this page's demo makes concrete) checkerboard poses that don't actually constrain the problem.
That last failure mode is common and easy to walk into by accident: if every pose is coplanar with the others in a way that doesn't vary enough — the board is always held roughly fronto-parallel to the camera, or every pose is a rotation about the same axis — the linear system for ω becomes rank-deficient, exactly as the degenerate preset in the demo above shows. The fix that every calibration guide repeats ("tilt the board, cover the whole frame, rotate it around different axes between shots") isn't superstition — it's making sure the v rows stacked into V actually span enough directions for Vb = 0 to pin down a unique b. More poses without more orientation diversity doesn't help; three well-chosen tilts beat thirty photos of the board held flat.
Cheat sheet
Recap
| Object | Symbol / form | Definition | Where it matters |
|---|---|---|---|
| Image of the absolute conic | ω = K⁻ᵀK⁻¹ | Symmetric 3×3, 5 DOF up to scale; same for every pose of one camera | The thing Zhang's method actually solves for |
| K from ω | ω = LLᵀ, K = (Lᵀ)⁻¹ | Cholesky decomposition, normalized so K₃₃ = 1 | Closed-form step after the linear solve |
| Zhang constraint 1 | h₁ᵀωh₂ = 0 | From r₁ᵀr₂ = 0 (rotation columns orthogonal) | One row of V per pose |
| Zhang constraint 2 | h₁ᵀωh₁ = h₂ᵀωh₂ | From r₁ᵀr₁ = r₂ᵀr₂ (rotation columns unit length) | Second row of V per pose |
| Minimum poses | ≥ 3, orientation-diverse | 2 constraints/pose, 5 unknowns up to scale ⟹ ≥2.5 poses | Fewer, or all-coplanar, ⟹ V rank-deficient |
| Full intrinsics | K = [[fx,s,cx],[0,fy,cy],[0,0,1]] | 5 DOF: focal lengths, skew, principal point | Part 4's demo fixed fx=fy, s=0 |
| Reprojection error | RMS, in pixels | Distance between reprojected and detected corners | < 0.5 px good, > 1 px investigate |