Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

The question

How should a belief and a measurement be combined?

Statistics usually hands you a sample and asks for a parameter. Bayesian inference hands you two distributions and asks for a third: start from a prior over an unknown, multiply by the likelihood of the data, normalise, and read off the posterior. For an arbitrary prior and an arbitrary likelihood that multiplication can be an intractable integral. But there is a special case where everything stays closed-form, and it happens to be the case that runs the machinery of robotics, tracking, and sensor fusion.

Take a single unknown number μ. Before seeing anything, you describe it with a Gaussian N(μ0, σ02). An observation arrives as a number x, and the sensor's error is Gaussian with variance σ2, so the likelihood of x given μ is a Gaussian centred at μ. Equivalently, the observation itself is a Gaussian centred at x. Bayes' rule multiplies the two curves point by point, then rescales so the area is one.

A Gaussian times a Gaussian is a Gaussian. The posterior is therefore again in the family you started from: a Gaussian. This self-reproducing property is called conjugacy, and it is what makes the arithmetic on this page possible without a single numerical integral.

💡 By the end of this part you'll see why the posterior precision is the sum of the prior and observation precisions, why the more precise source dominates the blend, and why doing this once and then again is exactly the recursive update a Kalman filter runs forever.
2

Same family in, same family out

Why the posterior stays Gaussian

Conjugacy is a promise about the shape of the answer. If your prior and your likelihood are both Gaussian, the posterior is Gaussian; you never leave the family. The prior contributes a curve that falls off like exp(−(μ−μ0)2/2σ02), the likelihood contributes one that falls off like exp(−(x−μ)2/2σ2), and multiplying adds their exponents. The result is a quadratic in μ, and a quadratic exponent is precisely a Gaussian. Completing the square gives both the new centre and the new width at once:

$$\mu_{\text{post}} = \frac{\mu_0 \tau_0 + x \tau}{\tau_0 + \tau}, \qquad \tau_{\text{post}} = \tau_0 + \tau, \qquad \tau_0 = \frac{1}{\sigma_0^2}, \quad \tau = \frac{1}{\sigma^2}$$

Read the second expression first, because it is the clean one: the posterior precision is the sum of the prior precision and the observation precision. The first expression then says the posterior mean is the average of the two centres, weighted by those same precisions. There is no approximation here; this is the exact posterior under the stated model.

Conjugacy is a convenience, not a law of nature. Change the likelihood to something non-Gaussian, or put a hard constraint on the parameter, and the multiplication no longer closes; you are pushed toward the numerical methods of the MCMC part of this series. The Gaussian–Gaussian pair is the one case where the algebra is simple enough to reason about by eye, and it happens to be the case that matters most in practice. It is the same algebra that appears whenever independent Gaussian errors are added, as in the probability guide: adding variances there is adding reciprocals of precisions here.

The formula also makes clear what the prior is for. It is not a mystical extra ingredient; it is one more Gaussian observation, carrying its own precision. A vague prior carries little precision and is quickly overwhelmed. A confident prior carries a lot and resists a noisy measurement. Everything in the next section follows from that single reading.

3

Precision is the currency

Why the sharper source wins

Rewrite the update in terms of precision, τ = 1/σ2. Variance is the error you carry; precision is the information you have. The update says the posteriors add their information, and the mean becomes a weighted average whose weights are the precisions themselves:

$$\mu_{\text{post}} = \frac{\tau_0}{\tau_0 + \tau}\,\mu_0 + \frac{\tau}{\tau_0 + \tau}\,x$$

Two consequences follow immediately, and both are visible in the demo below. First, the answer is always sharper than either input: precisions are positive, so the sum exceeds each one, and the posterior variance is smaller than both σ02 and σ2. Fusing independent information can never make you less certain. Second, the blend leans toward whichever source has the larger precision. If the sensor is ten times more precise than your prior, it receives roughly ten elevenths of the weight.

Push the limits and the formula becomes a statement about humility. Let the observation variance go to infinity — a useless sensor — and τ → 0; the posterior collapses onto the prior. Let the prior variance go to infinity — you knew nothing — and τ0 → 0; the posterior collapses onto the observation. A flat prior is the formal way of saying "I have no opinion", and with it the Gaussian update reduces to "believe the data".

Independence is doing real work in the word "add". If two readings share a calibration error, adding their precisions double-counts the same information and you will be overconfident. That failure, not the arithmetic, is what breaks sensor fusion in practice, and it is the reason an engineer spends more time on error models than on the update equation.

4

Fuse two Gaussians

The signature interaction

The panel draws a prior (blue) and an observation (pink), both Gaussian. Drag either handle along the axis to move its mean; use the sliders to change each width. Press Merge and the posterior (dark) animates in, drawn from the same formula the text derived. Watch the posterior: it sits between the means, nearer the sharper curve, and it is always narrower than both.

Try dragging the observation far from the prior. If the observation is narrow, the posterior follows it almost completely; widen the observation and the prior pulls the answer back. Set the two widths equal and the posterior lands exactly halfway between the means, because equal variance means equal precision. Every movement recomputes the posterior from the precision formula; nothing is cached and nothing is approximated.

Prior, observation, and the posterior they fuse into. Drag the two handles to move the means; the sliders set the widths.

The readout also computes the same update the Kalman filter will use later in this series. In one dimension, with H = 1 and measurement noise R = σ2, the filter forms the gain K = σ02/(σ02+σ2), then μ′ = μ0 + K(x − μ0) and σ′2 = (1−K)σ02. Substitute K, simplify once, and you get the precision formula exactly. The page runs the Kalman update on your current numbers, and the readout confirms it lands on the same mean and variance to machine precision. The scalar Kalman update is not an analogy for Gaussian fusion; it is the same equation rearranged.

⚠ Fusion sharpens, but only under the model. The posterior variance is always below both inputs, which looks like a free lunch. It is not. That guarantee is conditional on the prior mean, the observation mean and both variances being what you say they are. A sensor with a bias violates the model silently, and the fused estimate becomes confidently wrong.
5

Updating again is just updating

From one fusion to a recursive estimator

One fusion is an event. The interesting behaviour starts when you repeat it. The posterior is a Gaussian, so it can play the role of the prior in the next fusion. Feed in a second observation and its precision is added to the posterior precision; a third adds again. After n independent observations of the same quantity, each with precision τ, the total is τ0 + nτ and the variance falls like 1/(τ0 + nτ). The mean drifts toward the data, more and more slowly, as the accumulated precision grows.

That is a recursive estimator: a running summary — here just a mean and a variance — that you update in place as each measurement arrives, without storing the measurements themselves. The prior is not a fixed thing; it is whatever you believed a moment ago. This is why the Gaussian pair is the natural language of tracking. The full Kalman filter adds a prediction step between updates: it lets the state move, which inflates the variance, and then the measurement update sharpens it again. Predict, update, predict, update — the Gaussian fusion on this page, run forever, with a motion model stretched between the steps.

Sequential and batch updating agree. Fuse observations one at a time or all at once and the final posterior is identical, because precision addition does not care about order. That equivalence is what lets an online filter produce exactly the answer a statistician would get by keeping every measurement and fitting at the end. It is also what makes the filter computationally cheap: each step costs the same, regardless of how long the system has been running. When the hidden state is a full vector with a covariance matrix, the scalar reciprocal becomes a matrix inverse and the precision sum becomes a matrix sum, but the shape of the idea — add information, reweight, sharpen — is unchanged.

6

Where this shows up

One update, two worlds

Robotics

Wheel odometry meets a sensor

A robot has a motion model that predicts where it should be and a wheel encoder that reports how far it thinks it went; both are uncertain, and both are summarised by Gaussians. The fusion runs exactly like this: the model supplies a prior, the encoder supplies a measurement, and the precision-weighted average is the position estimate. Repeat it every timestep and the recursion above becomes a localization filter.

Vision & Geometry

Beliefs in a SLAM system

A SLAM backend holds a Gaussian belief over poses and landmarks and folds in every camera observation as a measurement. The SLAM chapter replaces the single unknown here with a large vector and the scalar precision with an information matrix, but the operation on each new measurement is still the fusion on this page — a weighted merge in which sharper constraints dominate the result.

The pattern is not particular to robots. Any pipeline that maintains a running estimate and folds in a stream of noisy measurements is doing Gaussian fusion somewhere: inertial navigation, satellite positioning, radar tracking, camera calibration, and the state estimate inside a visual odometry front end. The precision-weighted average, written once, is the whole idea.

Further reading

The references below treat conjugate updates as the analytic core of Bayesian inference rather than a special trick. If you take away one thing, take away the picture of two precision bars stacking into a taller one while the mean settles between them.

Cheat sheet

TermMeaning here
ConjugacyPrior and posterior lie in the same family, so the update stays closed-form
Precisionτ = 1/σ2; the information a Gaussian carries
FusionMultiplying prior and likelihood; precisions add, means blend
Posterior mean(μ0τ0 + xτ)/(τ0+τ) — a precision-weighted average
Posterior precisionτ0 + τ; always larger than either input, so the posterior is always sharper
DominanceThe source with larger precision receives the larger weight
Recursive updateYesterday's posterior becomes today's prior; variance falls like 1/(τ0+nτ)
Kalman linkWith H=1, the scalar Kalman update is this formula rearranged
7

Check your understanding

0/4 answered