Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

How to use this card

Filter, then follow the link

Every row of the first table links its idea to the part that introduces it, in the order the volume teaches them. The second table collects the named distributions; the conjugate-prior column is the bridge to Bayesian inference in the companion volume, where those pairs are the reason closed-form posteriors exist at all. Nothing here restates a proof — use it to find the part, then read the part.

No row matches that — try a looser word, or clear the box.

2

Terms, in the order they appear

From the part that builds it

TermMeaningIntroduced in
Sample space, outcome, event$\Omega$ is the set of all outcomes; an event is a subset of it.Part 1
Axioms of probabilityNon-negativity, normalisation, countable additivity.Part 1
Relative frequencyS_n/n, the proportion of successes in n trials.Part 1
Odds, log-odds, coherencep/(1-p) and its log; incoherent beliefs admit a Dutch book.Part 1
Permutation, combination, binomial coefficientn!/(n-k)!, $\binom{n}{k}=n!/(k!(n-k)!)$.Part 2
Birthday problem, collision probabilityChance of a repeated value when drawing from a finite set.Part 2
Conditional probability$P(A\mid B)=P(A\cap B)/P(B)$.Part 3
Multiplication / chain rule$P(A\cap B)=P(B)P(A\mid B)$ and its extension.Part 3
Law of total probabilitySumming over a partition of the sample space.Part 3
Bayes' rule$P(D\mid +)\propto P(+\mid D)P(D)$; posterior, prior, likelihood.Part 4
Base-rate fallacy, likelihood ratioThe prior dominates a rare-event test; LR multiplies the odds.Part 4
Independence, conditional independence$P(A\cap B)=P(A)P(B)$; the conditional version can hold when the plain one fails.Part 5
Explaining-away, Simpson's paradoxA collider induces dependence; pooling can reverse an association.Part 5
Random variable, pmf, cdfA function from $\Omega$ to $\mathbb{R}$; $F(x)=P(X\le x)$.Part 6
Expectation$\mathbb{E}[X]=\sum_k k\,P(X=k)$; linear whatever the dependence.Part 7
Variance, standard deviation$\operatorname{Var}(X)=\mathbb{E}[X^2]-\mathbb{E}[X]^2$.Part 8
Markov, Chebyshev, HoeffdingTail bounds that need only moments or a bounded range.Part 8
Discrete familiesBernoulli, binomial, geometric, negative binomial, Poisson.Part 9
Poisson limitBinomial with $np\to\lambda$ converges to Poisson.Part 9
Probability density$P(a\le X\le b)=\int_a^b f(x)\,dx$; f is not a probability.Part 10
Continuous familiesUniform, exponential, gamma, beta, Student-t, Gaussian.Part 11
Max-entropy characterisationThe Gaussian is the maximum-entropy density for a given variance.Part 11
Change of variables, inverse-CDF samplingDensities rescale by |g'|; F^{-1}(U) samples any distribution.Part 12
Joint, marginal, conditionalp(x,y), its row and column sums, and $p(y\mid x)$.Part 13
Covariance, correlation$\operatorname{Cov}(X,Y)$; $\rho$ is the normalised version.Part 14
Covariance matrix $\Sigma$Symmetric positive semi-definite; its eigenvectors are the ellipse axes.Part 14
Multivariate Gaussian$\mathcal{N}(\mu,\Sigma)$; conditioning, marginals, Mahalanobis distance.Part 15
Whitening$\Sigma^{-1/2}(x-\mu)$ makes the components independent standard normals.Part 15
Convolution, MGFThe density of a sum is a convolution; mgfs turn it into a product.Part 16
Law of large numbers$\bar X_n\to\mu$ in probability and almost surely.Part 17
Modes of convergenceIn probability, almost sure, in mean square, in distribution.Part 17
Central limit theorem$\sqrt{n}(\bar X_n-\mu)/\sigma\Rightarrow\mathcal{N}(0,1)$.Part 18
Berry–Esseen boundThe $O(1/\sqrt n)$ rate of CLT convergence.Part 18
Heavy tails, concentration of measureWhere the CLT fails; why high dimensions are counter-intuitive.Part 19
σ-algebra, measurable setThe events a probability is allowed to assign; Vitali sets are not measurable.Part 20
Measure, filtration, Borel–CantelliA probability is a measure with total mass one; filtrations encode information over time.Part 20
3

The distribution reference card

pmf or pdf, mean, variance, mgf, conjugate prior

Familypmf / pdfMeanVarianceMGFConjugate prior for
Bernoulli (p)p^k(1-p)^{1-k}pp(1-p)1-p+pe^tBeta (itself)
Binomial (n,p)$\binom{n}{k}p^k(1-p)^{n-k}$npnp(1-p)(1-p+pe^t)^nBeta
Geometric (p), on $1,2,\dots$(1-p)^{k-1}p1/p(1-p)/p^2$\dfrac{pe^t}{1-(1-p)e^t}$Beta
Negative binomial (r,p)$\binom{k-1}{r-1}p^r(1-p)^{k-r}$r/pr(1-p)/p^2$\left(\dfrac{p}{1-(1-p)e^t}\right)^r$Beta
Poisson $(\lambda)$$e^{-\lambda}\lambda^k/k!$$\lambda$$\lambda$$\exp(\lambda(e^t-1))$Gamma (rate)
Uniform (a,b)$\dfrac{1}{b-a}$$\dfrac{a+b}{2}$$\dfrac{(b-a)^2}{12}$$\dfrac{e^{tb}-e^{ta}}{t(b-a)}$—
Exponential $(\lambda)$$\lambda e^{-\lambda x}$$1/\lambda$$1/\lambda^2$$\dfrac{\lambda}{\lambda-t}\;(t<\lambda)$Gamma
Normal $(\mu,\sigma^2)$$\dfrac{1}{\sigma\sqrt{2\pi}}e^{-(x-\mu)^2/2\sigma^2}$$\mu$$\sigma^2$$\exp(\mu t+\tfrac12\sigma^2t^2)$Normal (mean), inverse-gamma (variance)
Gamma $(k,\theta)$$\dfrac{x^{k-1}e^{-x/\theta}}{\theta^k\Gamma(k)}$$k\theta$$k\theta^2$$(1-\theta t)^{-k}\;(t<1/\theta)$Gamma (Poisson rate), inverse-gamma (normal variance)
Beta (a,b)$\dfrac{x^{a-1}(1-x)^{b-1}}{B(a,b)}$$\dfrac{a}{a+b}$$\dfrac{ab}{(a+b)^2(a+b+1)}$no simple closed formBinomial / Bernoulli
Student-t $(\nu)$$\dfrac{\Gamma((\nu+1)/2)}{\sqrt{\nu\pi}\,\Gamma(\nu/2)}\left(1+\tfrac{x^2}{\nu}\right)^{-(\nu+1)/2}$$0\;(\nu>1)$$\dfrac{\nu}{\nu-2}\;(\nu>2)$undefined—
Chi-squared (k)$\dfrac{x^{k/2-1}e^{-x/2}}{2^{k/2}\Gamma(k/2)}$k2k$(1-2t)^{-k/2}\;(t<1/2)$—
Cauchy $(x_0,\gamma)$$\dfrac{1}{\pi\gamma\left[1+\left(\frac{x-x_0}{\gamma}\right)^2\right]}$undefinedundefinedundefined—
Lognormal $(\mu,\sigma^2)$$\dfrac{1}{x\sigma\sqrt{2\pi}}e^{-(\ln x-\mu)^2/2\sigma^2}$$e^{\mu+\sigma^2/2}$$(e^{\sigma^2}-1)e^{2\mu+\sigma^2}$undefined—

Quantiles and cdfs are computed by Prob.dist in assets/js/prob-viz.js; the regularised incomplete gamma and beta functions behind the gamma, chi-squared, Poisson, binomial, beta and Student-t cdfs are in Prob.sf. The conjugate-prior column is the reason a Bayesian update stays in closed form — the posterior lands back in the same family as the prior.

4

Where to go next

From foundations to algorithms

This volume ends at the limit theorems. The algorithms those theorems justify — maximum likelihood, intervals, sampling, Markov chains, MCMC, and the Bayes, Kalman and particle filters — are the companion volume, Probability in Action, Interactively. Its own routing table maps each of those ideas to the AI, vision and robotics page that consumes it.