Glossary, distributions and cheat cards
This is the series' back matter: the notation fixed in Part 1, the distribution and test tables the later acts lean on, the estimator and interval card, and a filterable index that links every term to the part that introduces it. When a symbol in a page is unfamiliar, find it here first. Probability — sample spaces, conditional probability, random variables, expectation and the named distributions as objects — is a prerequisite and is not defined here.
Notation, fixed once
The conventions every part follows
A random variable is a capital letter, X, and its observed value a lowercase one, x; the i-th observation is xi and the sample of n values is x1, …, xn, drawn i.i.d. An unknown parameter is θ and an estimator of it is θ̂ (a random variable before the data arrive, a number afterwards). The population mean and variance are μ and σ²; the sample mean and variance are x̄ and s². A probability density or mass is p(x | θ), the likelihood is the same function read as L(θ) = p(D | θ), and its logarithm is ℓ(θ).
Symbols
E[X] · Var(X) · Cov(X,Y) · ρ
θ̂ · SE(θ̂) · MSE · I(θ) · CRLB
α · β · power · p · R²
Reading order
Acts I–II are data and estimation, III is uncertainty and testing, IV is Bayesian inference, V is the statistics of learning, VI is estimation over time, and VII is causality. Every core part ends with a "Where this shows up" panel linking into the site's other guides.
The distribution table
Support, moments, and the prior that pairs with each
| Distribution | Support | Mean | Variance | Conjugate prior / role |
|---|---|---|---|---|
| Bernoulli(p) | {0, 1} | p | p(1−p) | Beta prior on p; one coin flip. |
| Binomial(n, p) | {0, …, n} | np | np(1−p) | Beta prior; heads in n flips. |
| Poisson(λ) | {0, 1, 2, …} | λ | λ | Gamma prior on λ; event counts. |
| Uniform(a, b) | [a, b] | (a+b)/2 | (b−a)²/12 | No conjugate; the "flat prior" of ignorance. |
| Normal(μ, σ²) | ℝ | μ | σ² | Normal–Normal; the Kalman filter's currency. |
| Exponential(λ) | [0, ∞) | 1/λ | 1/λ² | Gamma prior; waiting times. |
| Gamma(k, θ) | (0, ∞) | kθ | kθ² | Conjugate for Poisson and Exponential rates. |
| Beta(α, β) | [0, 1] | α/(α+β) | αβ / ((α+β)²(α+β+1)) | Conjugate for Bernoulli/Binomial; a distribution over probabilities. |
| Student tν | ℝ | 0 for ν>1 | ν/(ν−2) for ν>2 | Exact null for the t-test; fat tails. |
| χ²(k) | [0, ∞) | k | 2k | Null for variance tests and likelihood ratios. |
| F(d1, d2) | [0, ∞) | d2/(d2−2) | see text | Null for comparing variances and nested models. |
| Categorical(p1..k) | {1, …, k} | — | — | Dirichlet prior; the LLM sampling distribution. |
Which test?
From the question to the statistic and its null
The whole zoo is one idea: pick a statistic, work out its distribution when the null is true, and measure how far into the tail the data land. The table is the picker; Part 15 is the interactive version.
| Question | Test | Null distribution | Key assumption | Part |
|---|---|---|---|---|
| Is one mean different from a value? | one-sample t (or z if σ known) | tn−1 / Normal | roughly normal data, or a large n | 15 |
| Do two independent groups differ? | Student or Welch t | t with pooled or Welch df | independence; equal variance for pooled | 15 |
| Before/after the same units? | paired t | tn−1 | the paired differences are normal | 11 |
| Is a proportion different? | z / exact binomial | Normal / Binomial | independent trials, fixed n | 16 |
| Are two categorical variables associated? | χ² independence | χ²(r−1)(c−1) | expected counts not too small | 15 |
| Do three or more means differ? | F (ANOVA) | Fk−1, n−k | independent, normal, equal variance | 15 |
| Do two variances differ? | F test | Fn₁−1, n₂−1 | normal data | 15 |
| Is a bigger model worth it? | likelihood-ratio | χ²Δk | nested models, large n | 15 |
| No distributional assumptions? | permutation | empirical (shuffled labels) | exchangeability under the null | 12 |
| Many tests at once? | Bonferroni / Benjamini–Hochberg | — | independent or PRDS for FDR | 14 |
Estimator and interval card
The formulas the series builds, in one place
Sample mean, the unbiased sample variance, and the standard error of the mean. The √n is the whole reason more data helps.
Confidence intervals for a mean and a proportion. The interval is the random object; the parameter is fixed. Coverage, not probability, is what the 1−α names.
Maximum likelihood and the Cramér–Rao floor. The curvature of ℓ at its peak is the error bar; the MLE attains the floor asymptotically.
The error budget split into the part more data cannot fix and the part it can, and the Bayesian point estimate as a penalised likelihood.
Ordinary least squares and the covariance of its coefficients; the confidence band in Part 10 is the square root of the diagonal.
Bayes' rule, and Gaussian fusion as precision-weighted averaging — the scalar Kalman update of Part 17.
Glossary
Every term, linked to the part that introduces it
| Term | What it means | Introduced in |
|---|---|---|
| A | ||
| A/B test | A randomized comparison of two variants; randomization is what licenses the causal reading. | Part 30 |
| Alternative hypothesis | The claim a test is trying to detect; its distance from the null is the effect size. | Part 12 |
| ATE | Average treatment effect: the mean difference in potential outcomes across the population. | Part 32 |
| Autocorrelation | Correlation between successive MCMC samples; it inflates the variance of any estimate from the chain. | Part 19 |
| B | ||
| Bayes' rule | p(θ|D) p(D|θ)p(θ): posterior is likelihood times prior, up to a constant. | Part 16 |
| Bayes filter | The recursion that alternates a motion update (blur) with a measurement update (multiply) on a belief. | Part 27 |
| Bias | E[θ̂] − θ: the systematic part of an estimator's error, untouched by sample size. | Part 6 |
| Bias–variance | The split of expected squared error into bias², variance and irreducible noise; the U-curve of model complexity. | Part 21 |
| BIC | Bayesian information criterion, k ln n − 2 ln L: a penalised likelihood with a stronger complexity penalty than AIC. | Part 23 |
| Bootstrap | Resampling the data with replacement to approximate the sampling distribution of any statistic. | Part 9 |
| Brier score | Mean squared error of probability forecasts; a proper scoring rule for calibration. | Part 24 |
| C | ||
| Calibration | Agreement between predicted probabilities and observed frequencies; separate from ranking quality. | Part 24 |
| Central limit theorem | Averages of i.i.d. draws approach a normal shape; the standard error is σ/√n. | Part 5 |
| Collider | A node caused by two others; conditioning on it creates a spurious association between them. | Part 31 |
| Conditional distribution | The distribution of one variable given another; for a Gaussian its mean is linear in the conditioning value. | Part 26 |
| Confidence interval | A random interval with a stated long-run coverage; the parameter it covers is fixed. | Part 11 |
| Confounder | A common cause of treatment and outcome; it opens a back-door path that must be blocked. | Part 31 |
| Conjugate prior | A prior whose posterior stays in the same family as the likelihood — Beta–Binomial, Normal–Normal. | Part 16 |
| Consistency | An estimator that converges to the truth as n → ∞; unbiasedness is neither necessary nor sufficient. | Part 6 |
| Convolution | The distribution of a sum of independent variables; the Bayes filter's motion update is one. | Part 5 |
| Correlation | Standardised covariance in [−1, 1]; it measures linear association only. | Part 3 |
| Covariance | The average product of deviations; scale-dependent, and the off-diagonal of the covariance matrix. | Part 3 |
| Coverage | The fraction of repeated confidence intervals that contain the parameter; the meaning of "95%". | Part 11 |
| Credible interval | A range with a stated posterior probability; the Bayesian answer to an interval. | Part 18 |
| Cramér–Rao bound | A floor of 1/(nI(θ)) on the variance of any unbiased estimator. | Part 8 |
| Cross-validation | Rotating which fold is held out to estimate generalisation error without wasting data. | Part 23 |
| D | ||
| DAG | Directed acyclic graph: nodes are variables and arrows encode assumed direct causes. | Part 31 |
| Decision theory | Choosing an action to minimise expected loss; the Bayesian report is a decision, not a given. | Part 18 |
| Deprivation | Particle-filter collapse where all surviving particles occupy one region, so no alternative hypothesis remains. | Part 29 |
| Difference-in-differences | Comparing the change over time in a treated group with the change in a control group; needs parallel trends. | Part 32 |
| E | ||
| ECDF | Empirical cumulative distribution function: the fraction of data at or below each value. | Part 2 |
| Effective sample size | 1/Σw²: how many particles are effectively contributing; the resampling trigger. | Part 29 |
| Efficiency | How close an estimator's variance comes to the Cramér–Rao floor. | Part 8 |
| EKF | Extended Kalman filter: the Kalman update applied to a Taylor-linearised nonlinear model. | Part 28 |
| Estimator | A rule for turning a sample into a guess about a parameter; a random variable before the data arrive. | Part 6 |
| F | ||
| F distribution | The distribution of a ratio of variances; the null for ANOVA and nested-model tests. | Part 15 |
| False discovery rate | The expected fraction of rejections that are false; controlled by Benjamini–Hochberg. | Part 14 |
| Fisher information | Variance of the score, and the curvature of the log-likelihood; its reciprocal bounds the error. | Part 8 |
| Family-wise error rate | The probability of at least one false rejection across a family of tests; controlled by Bonferroni. | Part 14 |
| G | ||
| Gaussian | The normal distribution; its two parameters are mean and variance, and it is closed under adding and conditioning. | Part 3 |
| Gaussian fusion | Combining two Gaussian beliefs into a precision-weighted average; the scalar Kalman update. | Part 17 |
| Gibbs sampling | MCMC that cycles through coordinates, sampling each from its conditional given the rest. | Part 19 |
| H | ||
| Histogram | Counts per bin; its shape depends on bin width, a bias–variance choice. | Part 2 |
| Histogram filter | A Bayes filter whose belief is a discrete grid of probabilities. | Part 27 |
| HPD interval | Highest posterior density: the shortest range covering a given posterior mass. | Part 18 |
| I | ||
| i.i.d. | Independent and identically distributed; the assumption nearly every method in the series starts from. | Part 1 |
| Ignorability | Treatment is independent of potential outcomes given covariates; the assumption that makes adjustment work. | Part 32 |
| IPW | Inverse-probability weighting: reweight units by 1/p(T|X) to emulate a randomized trial. | Part 32 |
| IQR | Interquartile range, the span of the middle half; the basis of a robust box plot. | Part 2 |
| K | ||
| Kalman filter | Exact Bayesian filtering for linear-Gaussian systems; predict inflates, update fuses. | Part 28 |
| KDE | Kernel density estimate: a smoothed histogram whose bandwidth is the smoothing knob. | Part 2 |
| L | ||
| Laplace approximation | A Gaussian centred at the posterior mode with covariance equal to the inverse Hessian. | Part 20 |
| Lasso | L1-penalised regression, equivalent to a Laplace prior; it drives coefficients exactly to zero. | Part 22 |
| Likelihood | The probability of the data as a function of the parameters; not a probability over parameters. | Part 7 |
| Likelihood ratio | 2(ℓ₁−ℓ₀), distributed as χ² under the null; the general engine behind many tests. | Part 15 |
| Log-loss | Negative log-likelihood of probability forecasts; a proper scoring rule that punishes confident errors. | Part 24 |
| Loss function | The cost of an error; its choice decides whether the mean, median or mode is the report. | Part 18 |
| M | ||
| MAP | Maximum a posteriori: the mode of the posterior; a penalised maximum-likelihood estimate. | Part 22 |
| Marginal distribution | The distribution of a subset of variables, integrating the rest out. | Part 3 |
| Matching | Pairing treated units with similar controls to compare like with like. | Part 32 |
| MCMC | Markov chain Monte Carlo: sampling from a posterior by constructing a chain that has it as its stationary distribution. | Part 19 |
| MCL | Monte Carlo localisation: a particle filter over robot poses. | Part 29 |
| Mean squared error | E[(θ̂−θ)²] = bias² + variance: the full error budget. | Part 6 |
| Median | The middle value; robust to outliers where the mean is not. | Part 2 |
| Metropolis–Hastings | MCMC that proposes a move and accepts it with the density-ratio probability. | Part 19 |
| MLE | Maximum likelihood estimate: the parameter that makes the observed data most probable. | Part 7 |
| Multivariate Gaussian | The normal in n dimensions: a mean vector and a covariance matrix; marginals and conditionals are Gaussian. | Part 26 |
| N | ||
| Null distribution | What the test statistic would look like if the null were true; the p-value is a tail area of it. | Part 12 |
| Null hypothesis | The default claim, usually "no effect", that a test tries to falsify. | Part 12 |
| O | ||
| Optimism | The gap by which training error underestimates generalisation error; it grows with capacity. | Part 23 |
| Outlier | A value far from the rest; it pulls the mean but barely moves the median. | Part 2 |
| Overfitting | Fitting noise; low bias, high variance, poor generalisation. | Part 21 |
| P | ||
| p-hacking | Trying many analyses and reporting the significant one; it inflates the false-positive rate. | Part 14 |
| Parameter | A fixed, unknown number describing the population; the target of estimation. | Part 1 |
| Particle filter | A Bayes filter whose belief is a weighted cloud of samples, resampled when weights degenerate. | Part 29 |
| Permutation test | A test whose null is built by shuffling labels, making no distributional assumption. | Part 12 |
| Posterior | The updated belief p(θ|D): the whole answer, not just a point. | Part 16 |
| Power | The probability of detecting a real effect; 1 − β, rising with sample size and effect size. | Part 13 |
| Precision and recall | Of the flagged positives, how many were right; of the real positives, how many were found. | Part 24 |
| Predictive interval | A range for a future observation, wider than a confidence interval for a mean. | Part 25 |
| Prior | The belief before the data; the regulariser with a probabilistic reading. | Part 16 |
| Propensity score | p(T=1|X): the probability of treatment given covariates; the basis of matching and IPW. | Part 32 |
| p-value | The probability, under the null, of a statistic at least as extreme as observed; not the probability the null is true. | Part 12 |
| Q | ||
| Quantile | The value below which a given fraction of the data lies; the median is the 0.5 quantile. | Part 2 |
| R | ||
| Random variable | A function from outcomes to numbers; an estimate is one, built from a sample. | Part 1 |
| Randomization | Assigning treatment by chance so that covariates balance in expectation, identifying the causal effect. | Part 30 |
| Regression | Estimating a conditional mean as a linear function of predictors; least squares is its Gaussian MLE. | Part 10 |
| Regularization | Penalising complexity; equivalently, imposing a prior on the parameters. | Part 22 |
| Residual | Observed minus fitted; its pattern is the diagnostic for a regression. | Part 10 |
| Ridge | L2-penalised regression, equivalent to a Gaussian prior; it shrinks but rarely zeroes. | Part 22 |
| R-hat | A convergence diagnostic comparing between- and within-chain variance; near 1 is good. | Part 19 |
| ROC / AUC | The trade-off of true- and false-positive rates over all thresholds, and its area: a ranking measure. | Part 24 |
| S | ||
| Sample | The observed values; random, and the only thing you actually have. | Part 1 |
| Sampling distribution | The distribution of an estimator over repeated samples; its spread is the standard error. | Part 4 |
| Schur complement | The block formula for a conditional Gaussian's covariance, Σ₂₂ − Σ₂₁Σ₁₁⁻¹Σ₁₂. | Part 26 |
| Score function | The gradient of the log-likelihood; its variance is the Fisher information. | Part 8 |
| Sequential testing | Testing repeatedly as data arrive; naive peeking inflates the error rate, so alpha is spent instead. | Part 30 |
| Significance level | α: the false-rejection rate a test is willing to tolerate when the null holds. | Part 12 |
| Simpson's paradox | An association that reverses when data are pooled across a confounding group. | Part 31 |
| Skewness | Asymmetry of a distribution; positive skew puts the mean above the median. | Part 2 |
| Standard deviation | The square root of the variance; spread in the data's own units. | Part 2 |
| Standard error | The standard deviation of an estimator; σ/√n for a mean. | Part 5 |
| Systematic resampling | Drawing N equally spaced points from the weight CDF; low-variance and O(N). | Part 29 |
| T | ||
| t-test | A test of a mean or mean difference using the t null, which accounts for estimated variance. | Part 15 |
| Test statistic | A single number summarising the evidence; its null distribution sets the p-value. | Part 12 |
| Trace plot | The chain's value against iteration; a convergence diagnostic. | Part 19 |
| Type I / II error | False positive versus false negative; their rates are α and β, and power is 1−β. | Part 13 |
| U | ||
| Unbiased | An estimator whose average over repeated samples equals the parameter; not the same as accurate. | Part 6 |
| Uncertainty | Aleatoric (irreducible noise) versus epistemic (missing knowledge); the two behave differently under data. | Part 25 |
| V | ||
| Variance | The average squared deviation; its square root is the standard deviation. | Part 2 |
| Variational inference | Approximating a posterior by the member of a tractable family that maximises the ELBO. | Part 20 |
| W | ||
| Weight | A particle's importance weight, proportional to its measurement likelihood. | Part 29 |
| Welch's t | A two-sample t test that does not assume equal variances. | Part 15 |
| Z | ||
| z-score | A value standardised by its mean and standard deviation; the z-test's statistic. | Part 15 |