Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

Notation, fixed once

The conventions every part follows

A random variable is a capital letter, X, and its observed value a lowercase one, x; the i-th observation is xi and the sample of n values is x1, …, xn, drawn i.i.d. An unknown parameter is θ and an estimator of it is θ̂ (a random variable before the data arrive, a number afterwards). The population mean and variance are μ and σ²; the sample mean and variance are x̄ and s². A probability density or mass is p(x | θ), the likelihood is the same function read as L(θ) = p(D | θ), and its logarithm is ℓ(θ).

Symbols

E[X] · Var(X) · Cov(X,Y) · ρ
θ̂ · SE(θ̂) · MSE · I(θ) · CRLB
α · β · power · p · R²

Reading order

Acts I–II are data and estimation, III is uncertainty and testing, IV is Bayesian inference, V is the statistics of learning, VI is estimation over time, and VII is causality. Every core part ends with a "Where this shows up" panel linking into the site's other guides.

2

The distribution table

Support, moments, and the prior that pairs with each

DistributionSupportMeanVarianceConjugate prior / role
Bernoulli(p){0, 1}pp(1−p)Beta prior on p; one coin flip.
Binomial(n, p){0, …, n}npnp(1−p)Beta prior; heads in n flips.
Poisson(λ){0, 1, 2, …}λλGamma prior on λ; event counts.
Uniform(a, b)[a, b](a+b)/2(b−a)²/12No conjugate; the "flat prior" of ignorance.
Normal(μ, σ²)ℝμσ²Normal–Normal; the Kalman filter's currency.
Exponential(λ)[0, ∞)1/λ1/λ²Gamma prior; waiting times.
Gamma(k, θ)(0, ∞)kθkθ²Conjugate for Poisson and Exponential rates.
Beta(α, β)[0, 1]α/(α+β)αβ / ((α+β)²(α+β+1))Conjugate for Bernoulli/Binomial; a distribution over probabilities.
Student tνℝ0 for ν>1ν/(ν−2) for ν>2Exact null for the t-test; fat tails.
χ²(k)[0, ∞)k2kNull for variance tests and likelihood ratios.
F(d1, d2)[0, ∞)d2/(d2−2)see textNull for comparing variances and nested models.
Categorical(p1..k){1, …, k}——Dirichlet prior; the LLM sampling distribution.
3

Which test?

From the question to the statistic and its null

The whole zoo is one idea: pick a statistic, work out its distribution when the null is true, and measure how far into the tail the data land. The table is the picker; Part 15 is the interactive version.

QuestionTestNull distributionKey assumptionPart
Is one mean different from a value?one-sample t (or z if σ known)tn−1 / Normalroughly normal data, or a large n15
Do two independent groups differ?Student or Welch tt with pooled or Welch dfindependence; equal variance for pooled15
Before/after the same units?paired ttn−1the paired differences are normal11
Is a proportion different?z / exact binomialNormal / Binomialindependent trials, fixed n16
Are two categorical variables associated?χ² independenceχ²(r−1)(c−1)expected counts not too small15
Do three or more means differ?F (ANOVA)Fk−1, n−kindependent, normal, equal variance15
Do two variances differ?F testFn₁−1, n₂−1normal data15
Is a bigger model worth it?likelihood-ratioχ²Δknested models, large n15
No distributional assumptions?permutationempirical (shuffled labels)exchangeability under the null12
Many tests at once?Bonferroni / Benjamini–Hochberg—independent or PRDS for FDR14
4

Estimator and interval card

The formulas the series builds, in one place

$$\bar{x} = \frac{1}{n}\sum_{i=1}^n x_i, \qquad s^2 = \frac{1}{n-1}\sum_{i=1}^n (x_i-\bar{x})^2, \qquad \mathrm{SE}(\bar{x}) = \frac{s}{\sqrt{n}}$$

Sample mean, the unbiased sample variance, and the standard error of the mean. The √n is the whole reason more data helps.

$$\bar{x} \pm t_{n-1,\,1-\alpha/2}\,\frac{s}{\sqrt{n}}, \qquad \hat{p} \pm z_{1-\alpha/2}\sqrt{\frac{\hat{p}(1-\hat{p})}{n}}$$

Confidence intervals for a mean and a proportion. The interval is the random object; the parameter is fixed. Coverage, not probability, is what the 1−α names.

$$\hat{\theta}_{\text{MLE}} = \arg\max_\theta \ell(\theta), \qquad \mathrm{Var}(\hat{\theta}) \gtrsim \frac{1}{n\,I(\theta)}$$

Maximum likelihood and the Cramér–Rao floor. The curvature of ℓ at its peak is the error bar; the MLE attains the floor asymptotically.

$$\mathrm{MSE}(\hat{\theta}) = \mathrm{bias}^2 + \mathrm{Var}(\hat{\theta}), \qquad \hat{\theta}_{\text{MAP}} = \arg\max_\theta \big[\ell(\theta) + \log p(\theta)\big]$$

The error budget split into the part more data cannot fix and the part it can, and the Bayesian point estimate as a penalised likelihood.

$$\hat{\beta} = (X^\top X)^{-1}X^\top y, \qquad \mathrm{Var}(\hat{\beta}) = \sigma^2 (X^\top X)^{-1}$$

Ordinary least squares and the covariance of its coefficients; the confidence band in Part 10 is the square root of the diagonal.

$$p(\theta \mid D) = \frac{p(D \mid \theta)\,p(\theta)}{p(D)}, \qquad \mu_{\text{post}} = \frac{\mu_0/\sigma_0^2 + \bar{x}/\sigma^2}{1/\sigma_0^2 + 1/\sigma^2}$$

Bayes' rule, and Gaussian fusion as precision-weighted averaging — the scalar Kalman update of Part 17.

5

Glossary

Every term, linked to the part that introduces it

💡 Filter the list. Type any fragment — a term, a symbol, an idea — and the table hides non-matching rows and reports how many remain. Or narrow by category with the buttons. Matching is case-insensitive and looks at both the term and its definition.
TermWhat it meansIntroduced in
A
A/B testA randomized comparison of two variants; randomization is what licenses the causal reading.Part 30
Alternative hypothesisThe claim a test is trying to detect; its distance from the null is the effect size.Part 12
ATEAverage treatment effect: the mean difference in potential outcomes across the population.Part 32
AutocorrelationCorrelation between successive MCMC samples; it inflates the variance of any estimate from the chain.Part 19
B
Bayes' rulep(θ|D) p(D|θ)p(θ): posterior is likelihood times prior, up to a constant.Part 16
Bayes filterThe recursion that alternates a motion update (blur) with a measurement update (multiply) on a belief.Part 27
BiasE[θ̂] − θ: the systematic part of an estimator's error, untouched by sample size.Part 6
Bias–varianceThe split of expected squared error into bias², variance and irreducible noise; the U-curve of model complexity.Part 21
BICBayesian information criterion, k ln n − 2 ln L: a penalised likelihood with a stronger complexity penalty than AIC.Part 23
BootstrapResampling the data with replacement to approximate the sampling distribution of any statistic.Part 9
Brier scoreMean squared error of probability forecasts; a proper scoring rule for calibration.Part 24
C
CalibrationAgreement between predicted probabilities and observed frequencies; separate from ranking quality.Part 24
Central limit theoremAverages of i.i.d. draws approach a normal shape; the standard error is σ/√n.Part 5
ColliderA node caused by two others; conditioning on it creates a spurious association between them.Part 31
Conditional distributionThe distribution of one variable given another; for a Gaussian its mean is linear in the conditioning value.Part 26
Confidence intervalA random interval with a stated long-run coverage; the parameter it covers is fixed.Part 11
ConfounderA common cause of treatment and outcome; it opens a back-door path that must be blocked.Part 31
Conjugate priorA prior whose posterior stays in the same family as the likelihood — Beta–Binomial, Normal–Normal.Part 16
ConsistencyAn estimator that converges to the truth as n → ∞; unbiasedness is neither necessary nor sufficient.Part 6
ConvolutionThe distribution of a sum of independent variables; the Bayes filter's motion update is one.Part 5
CorrelationStandardised covariance in [−1, 1]; it measures linear association only.Part 3
CovarianceThe average product of deviations; scale-dependent, and the off-diagonal of the covariance matrix.Part 3
CoverageThe fraction of repeated confidence intervals that contain the parameter; the meaning of "95%".Part 11
Credible intervalA range with a stated posterior probability; the Bayesian answer to an interval.Part 18
Cramér–Rao boundA floor of 1/(nI(θ)) on the variance of any unbiased estimator.Part 8
Cross-validationRotating which fold is held out to estimate generalisation error without wasting data.Part 23
D
DAGDirected acyclic graph: nodes are variables and arrows encode assumed direct causes.Part 31
Decision theoryChoosing an action to minimise expected loss; the Bayesian report is a decision, not a given.Part 18
DeprivationParticle-filter collapse where all surviving particles occupy one region, so no alternative hypothesis remains.Part 29
Difference-in-differencesComparing the change over time in a treated group with the change in a control group; needs parallel trends.Part 32
E
ECDFEmpirical cumulative distribution function: the fraction of data at or below each value.Part 2
Effective sample size1/Σw²: how many particles are effectively contributing; the resampling trigger.Part 29
EfficiencyHow close an estimator's variance comes to the Cramér–Rao floor.Part 8
EKFExtended Kalman filter: the Kalman update applied to a Taylor-linearised nonlinear model.Part 28
EstimatorA rule for turning a sample into a guess about a parameter; a random variable before the data arrive.Part 6
F
F distributionThe distribution of a ratio of variances; the null for ANOVA and nested-model tests.Part 15
False discovery rateThe expected fraction of rejections that are false; controlled by Benjamini–Hochberg.Part 14
Fisher informationVariance of the score, and the curvature of the log-likelihood; its reciprocal bounds the error.Part 8
Family-wise error rateThe probability of at least one false rejection across a family of tests; controlled by Bonferroni.Part 14
G
GaussianThe normal distribution; its two parameters are mean and variance, and it is closed under adding and conditioning.Part 3
Gaussian fusionCombining two Gaussian beliefs into a precision-weighted average; the scalar Kalman update.Part 17
Gibbs samplingMCMC that cycles through coordinates, sampling each from its conditional given the rest.Part 19
H
HistogramCounts per bin; its shape depends on bin width, a bias–variance choice.Part 2
Histogram filterA Bayes filter whose belief is a discrete grid of probabilities.Part 27
HPD intervalHighest posterior density: the shortest range covering a given posterior mass.Part 18
I
i.i.d.Independent and identically distributed; the assumption nearly every method in the series starts from.Part 1
IgnorabilityTreatment is independent of potential outcomes given covariates; the assumption that makes adjustment work.Part 32
IPWInverse-probability weighting: reweight units by 1/p(T|X) to emulate a randomized trial.Part 32
IQRInterquartile range, the span of the middle half; the basis of a robust box plot.Part 2
K
Kalman filterExact Bayesian filtering for linear-Gaussian systems; predict inflates, update fuses.Part 28
KDEKernel density estimate: a smoothed histogram whose bandwidth is the smoothing knob.Part 2
L
Laplace approximationA Gaussian centred at the posterior mode with covariance equal to the inverse Hessian.Part 20
LassoL1-penalised regression, equivalent to a Laplace prior; it drives coefficients exactly to zero.Part 22
LikelihoodThe probability of the data as a function of the parameters; not a probability over parameters.Part 7
Likelihood ratio2(ℓ₁−ℓ₀), distributed as χ² under the null; the general engine behind many tests.Part 15
Log-lossNegative log-likelihood of probability forecasts; a proper scoring rule that punishes confident errors.Part 24
Loss functionThe cost of an error; its choice decides whether the mean, median or mode is the report.Part 18
M
MAPMaximum a posteriori: the mode of the posterior; a penalised maximum-likelihood estimate.Part 22
Marginal distributionThe distribution of a subset of variables, integrating the rest out.Part 3
MatchingPairing treated units with similar controls to compare like with like.Part 32
MCMCMarkov chain Monte Carlo: sampling from a posterior by constructing a chain that has it as its stationary distribution.Part 19
MCLMonte Carlo localisation: a particle filter over robot poses.Part 29
Mean squared errorE[(θ̂−θ)²] = bias² + variance: the full error budget.Part 6
MedianThe middle value; robust to outliers where the mean is not.Part 2
Metropolis–HastingsMCMC that proposes a move and accepts it with the density-ratio probability.Part 19
MLEMaximum likelihood estimate: the parameter that makes the observed data most probable.Part 7
Multivariate GaussianThe normal in n dimensions: a mean vector and a covariance matrix; marginals and conditionals are Gaussian.Part 26
N
Null distributionWhat the test statistic would look like if the null were true; the p-value is a tail area of it.Part 12
Null hypothesisThe default claim, usually "no effect", that a test tries to falsify.Part 12
O
OptimismThe gap by which training error underestimates generalisation error; it grows with capacity.Part 23
OutlierA value far from the rest; it pulls the mean but barely moves the median.Part 2
OverfittingFitting noise; low bias, high variance, poor generalisation.Part 21
P
p-hackingTrying many analyses and reporting the significant one; it inflates the false-positive rate.Part 14
ParameterA fixed, unknown number describing the population; the target of estimation.Part 1
Particle filterA Bayes filter whose belief is a weighted cloud of samples, resampled when weights degenerate.Part 29
Permutation testA test whose null is built by shuffling labels, making no distributional assumption.Part 12
PosteriorThe updated belief p(θ|D): the whole answer, not just a point.Part 16
PowerThe probability of detecting a real effect; 1 − β, rising with sample size and effect size.Part 13
Precision and recallOf the flagged positives, how many were right; of the real positives, how many were found.Part 24
Predictive intervalA range for a future observation, wider than a confidence interval for a mean.Part 25
PriorThe belief before the data; the regulariser with a probabilistic reading.Part 16
Propensity scorep(T=1|X): the probability of treatment given covariates; the basis of matching and IPW.Part 32
p-valueThe probability, under the null, of a statistic at least as extreme as observed; not the probability the null is true.Part 12
Q
QuantileThe value below which a given fraction of the data lies; the median is the 0.5 quantile.Part 2
R
Random variableA function from outcomes to numbers; an estimate is one, built from a sample.Part 1
RandomizationAssigning treatment by chance so that covariates balance in expectation, identifying the causal effect.Part 30
RegressionEstimating a conditional mean as a linear function of predictors; least squares is its Gaussian MLE.Part 10
RegularizationPenalising complexity; equivalently, imposing a prior on the parameters.Part 22
ResidualObserved minus fitted; its pattern is the diagnostic for a regression.Part 10
RidgeL2-penalised regression, equivalent to a Gaussian prior; it shrinks but rarely zeroes.Part 22
R-hatA convergence diagnostic comparing between- and within-chain variance; near 1 is good.Part 19
ROC / AUCThe trade-off of true- and false-positive rates over all thresholds, and its area: a ranking measure.Part 24
S
SampleThe observed values; random, and the only thing you actually have.Part 1
Sampling distributionThe distribution of an estimator over repeated samples; its spread is the standard error.Part 4
Schur complementThe block formula for a conditional Gaussian's covariance, Σ₂₂ − Σ₂₁Σ₁₁⁻¹Σ₁₂.Part 26
Score functionThe gradient of the log-likelihood; its variance is the Fisher information.Part 8
Sequential testingTesting repeatedly as data arrive; naive peeking inflates the error rate, so alpha is spent instead.Part 30
Significance levelα: the false-rejection rate a test is willing to tolerate when the null holds.Part 12
Simpson's paradoxAn association that reverses when data are pooled across a confounding group.Part 31
SkewnessAsymmetry of a distribution; positive skew puts the mean above the median.Part 2
Standard deviationThe square root of the variance; spread in the data's own units.Part 2
Standard errorThe standard deviation of an estimator; σ/√n for a mean.Part 5
Systematic resamplingDrawing N equally spaced points from the weight CDF; low-variance and O(N).Part 29
T
t-testA test of a mean or mean difference using the t null, which accounts for estimated variance.Part 15
Test statisticA single number summarising the evidence; its null distribution sets the p-value.Part 12
Trace plotThe chain's value against iteration; a convergence diagnostic.Part 19
Type I / II errorFalse positive versus false negative; their rates are α and β, and power is 1−β.Part 13
U
UnbiasedAn estimator whose average over repeated samples equals the parameter; not the same as accurate.Part 6
UncertaintyAleatoric (irreducible noise) versus epistemic (missing knowledge); the two behave differently under data.Part 25
V
VarianceThe average squared deviation; its square root is the standard deviation.Part 2
Variational inferenceApproximating a posterior by the member of a tractable family that maximises the ELBO.Part 20
W
WeightA particle's importance weight, proportional to its measurement likelihood.Part 29
Welch's tA two-sample t test that does not assume equal variances.Part 15
Z
z-scoreA value standardised by its mean and standard deviation; the z-test's statistic.Part 15