The continuous families and the Gaussian
When a random variable can take any value in an interval, the histogram from the last part becomes a density, and densities come in families. A handful of them cover almost everything you will meet: the flat one that admits total ignorance, the exponential that describes memoryless waiting, the gamma that adds waiting, the beta that lives on an interval, and the Student-t that trades a Gaussian's light tail for a heavy one. Behind all of them sits one curve that keeps reappearing — the Gaussian. It is not just a convenient formula. It is the distribution with the largest possible entropy for a given variance, and it is the shape that sums of almost anything are attracted to. This part builds the families, then explains why the bell is inevitable.
The question
A handful of shapes, and one that keeps winning
A density is a function that tells you how to weight an interval: the chance that a continuous variable lands between a and b is the area under the curve there. The previous part fixed what that means and why the height at a point is not itself a probability. This part asks the practical follow-up. Which curves should you actually reach for, and why do the same few keep turning up in statistics, physics, machine learning and robotics?
The answer is that each family encodes a specific structural fact about the world. The uniform distribution says the only thing you know is the range. The exponential distribution says the hazard of an event does not change with time, which forces a memoryless waiting time. The gamma distribution is what you get when several such waits pile up. The beta distribution is the natural shape for an unknown proportion confined to the unit interval. The Student-t distribution is what appears when you estimate a centre from a small sample and refuse to pretend the variance is known. And the Gaussian is what appears when many small independent influences accumulate.
That last sentence deserves a whole section, because the Gaussian arrives by two completely different roads. The first is the central limit theorem: add up many small independent contributions and the sum's shape relaxes to a bell, whatever the pieces looked like. The second is information theory. If you know only a mean and a variance and want the distribution that assumes nothing else, the one with the most entropy — the least committal, in an exact sense — is the Gaussian. Two unrelated principles name the same curve. That coincidence is the reason the bell is everywhere.
We will build a switchable panel for the six continuous families, watch exponential waiting forget the past, look at how the Student-t tail outruns the Gaussian's, and finish with an entropy experiment where smoothing a stubborn distribution toward a bell makes its entropy climb to the Gaussian ceiling. By the end you should be able to name a family from a description of the mechanism, write down its density, mean and variance, and explain why the Gaussian is special rather than merely common.
Uniform and exponential
Ignorance, and waiting that forgets
The uniform distribution on [a,b] is the density that says every sub-interval of the same length is equally likely. It is the honest encoding of a bounded range with no other information, and it is the distribution you sample from whenever a computer draws a number. Its mean sits at the midpoint and its variance grows with the square of the width, so a uniform on a wide interval has a large spread and a uniform on a narrow one is almost a sharp spike.
The exponential distribution is the waiting time until the next event when events arrive at a constant average rate $\lambda$. Picture radioactive decays, clicks of a Geiger counter, or the gap between two requests to a server. Its density is largest at zero and decays geometrically, and its mean is $1/\lambda$ — the higher the rate, the shorter the typical wait.
What makes the exponential special is memorylessness. Suppose you have already waited s seconds for the event and it has not happened. The chance you must wait another t seconds is exactly the chance a fresh clock would need t seconds, with no correction for the time already spent. Algebraically, the conditional tail equals the unconditional tail.
This is not a curiosity; it is a characterisation. Among the families on $[0,\infty)$ with a constant hazard, the exponential is the only memoryless one. If a process sometimes "wears out" or "gets safer" as it ages, its hazard changes and the exponential is the wrong model. If it does not — arrivals at a fixed rate, decays of stable atoms — the exponential is forced. The panel below is the quickest way to internalise all six families at once, and the demo after it shows the memorylessness identity as two shaded tails of the same area.
Choose a family, then move its two parameters. The solid curve is the pdf and the dashed staircase is the cdf; the axes carry both a density scale and a probability scale up to one.
Top: the unconditional exponential pdf, with the tail beyond t shaded. Bottom: the same law conditioned on X>s, which simply restarts at s; its tail beyond s+t has exactly the same area.
Notice how the two shaded areas always agree, whatever you do to s. The conditional distribution is not the original tail squeezed into a smaller window; it is a brand new exponential sitting at s, identical in shape to the one at zero. That is what "memoryless" means, and it is why the exponential is the continuous analogue of the geometric distribution from the discrete families.
Gamma and beta
Waiting for the kth event, and the shape of a proportion
Add up k independent exponential waits of rate $\lambda$ and you get the time of the k-th arrival in a Poisson process. That sum has the gamma distribution. Its two parameters are a shape k>0 and a scale $\theta$, with the exponential as the special case k=1. Because it is a sum, the mean and variance add: the mean is $k\theta$ and the variance is $k\theta^2$, so the spread grows with the square root of the shape while the mean grows linearly. For a large shape the gamma becomes a rounded, nearly symmetric hump; for a small shape it peaks sharply near zero and trails a long right tail.
The gamma is not only about waiting. It is the conjugate prior for the rate of a Poisson process, and its $\chi^2$ relative is the sampling distribution of a sum of squared Gaussians, which is why every variance test in statistics eventually reaches for it.
The beta distribution lives on the unit interval, which makes it the natural language for an unknown proportion or probability. Two positive parameters $\alpha$ and $\beta$ bend the density: with $\alpha=\beta=1$ it is flat, with both above one it becomes a rounded hump in the middle, and with both below one it becomes U-shaped, piling mass at zero and one. Read $\alpha$ and $\beta$ as pseudo-counts of successes and failures and the beta is exactly the posterior for a success probability after seeing $\alpha-1$ successes and $\beta-1$ failures. Its moments are simple ratios of the parameters, which is why it slots into Bayesian updates so cleanly.
Top: gamma densities with the current shape in colour and reference shapes dashed. Bottom: beta densities on the unit interval; the dashed line is the uniform reference.
Drag the gamma shape from below one to well above one and watch the density migrate from a spike at the origin to a bell further out. Then drag the beta parameters to opposite extremes and watch the mass crowd into one corner: $\alpha$ controls the pressure toward one, $\beta$ the pressure toward zero. Two parameters, two knobs, and an enormous range of shapes in between.
Student-t and heavy tails
What a small sample does to a bell
Suppose you draw a sample from a Gaussian whose variance you do not know and estimate the variance from the same sample. Standardising the sample mean by the estimated standard error no longer gives a Gaussian; it gives the Student-t distribution with $\nu=n-1$ degrees of freedom. Formally, take an independent standard normal Z and an independent chi-squared V with $\nu$ degrees of freedom, and set $T=Z/\sqrt{V/\nu}$. Because the denominator is itself random, the ratio sometimes takes large values, and those occasional large values fatten the tails of the density.
The shape of that tail is the point. A Gaussian tail decays like e^{-x^2/2}, which is faster than any power law, so extreme values essentially never happen. The Student-t tail decays like a power, $x^{-(\nu+1)}$, so a value four standard deviations out is not a miracle under small $\nu$ — it is merely unlikely. The smaller the degrees of freedom, the heavier the tail; at $\nu=1$ the distribution is Cauchy and is so heavy that its mean fails to exist in the usual sense, and at $\nu=2$ the mean exists but the variance does not. As $\nu\to\infty$ the fattening washes out and the density converges to the standard normal.
This is why robust statistics uses the t distribution. If your data are contaminated by occasional outliers — a mislabelled example, a sensor glitch, a fat-fingered entry — a likelihood built on the Gaussian will let that one point drag the fitted parameters a long way, because it assigns such points an astronomically small probability. The t likelihood assigns them a merely small probability and refuses to be shocked. The demo below overlays the two tails on a log scale so the difference is visible rather than asserted: near the centre the curves are nearly identical, and far out they separate by orders of magnitude.
Survival probabilities P(X>x) on a base-ten log scale. The dashed curve is the standard normal; the solid curve is the Student-t. Lower $\nu$ means a heavier tail.
The Gaussian, derived
Two roads to the same bell
The normal density is the familiar exponential of a negative square, normalised so that its total area is one. The constant $\sqrt{2\pi}$ is not arbitrary; it is exactly the value that makes the Gaussian integral $\int e^{-x^2/2}\,dx=\sqrt{2\pi}$ come out right, a fact usually proved by squaring the integral and switching to polar coordinates.
The central limit theorem. Take independent copies of almost any distribution with finite mean $\mu$ and finite variance $\sigma^2$. Standardise their average by subtracting $\mu$ and dividing by $\sigma/\sqrt{n}$, and the distribution of that quantity converges to the standard normal as n grows. The base distribution is irrelevant beyond its first two moments. This is why measurement errors, aggregate demand, the height of a population and the sum of many small perturbations all look Gaussian: each is a sum, and sums relax to the bell. The central limit theorem part turns this into a live simulation where you choose the base distribution and slide the sample size.
The maximum-entropy derivation. The second road starts from a different question. Among all densities on the line with a prescribed mean $\mu$ and variance $\sigma^2$, which one is the least committal? Measure commitment by differential entropy, $h(f)=-\int f\log f$, and maximise it subject to the two moment constraints using Lagrange multipliers. The variational condition forces $\log f(x)$ to be a quadratic in x, and the normalisation, mean and variance conditions pin the coefficients. The unique maximiser is the Gaussian, and its entropy has a clean closed form.
The two roads meet at the same curve. One says the Gaussian is what sums become; the other says it is what you should assume when you know only two moments and nothing else. The practical reading is that a Gaussian assumption is a maximal-ignorance assumption, and any departure from it — skewness, heavy tails, bounded support — is extra information that a Gaussian throws away. The last demo makes the entropy road visible: it starts with a bimodal distribution standardised to mean zero and variance $\sigma^2$, then blends it toward $N(0,\sigma^2)$ with increasing weight. The entropy rises monotonically to the Gaussian ceiling, because the blend is climbing toward the least-committal shape with that variance.
Top: the fixed-variance start, the Gaussian target, and the current blend. Bottom: differential entropy against the blend weight, with the Gaussian ceiling marked. Both distributions share a mean of zero and a variance of one, so only their shape changes.
There is a delicate point hiding in the ceiling. Differential entropy is not invariant under a change of units, which is why the ceiling carries the $\sigma^2$ and why the comparison is meaningful only at fixed variance. A narrow Gaussian has lower entropy than a wide one, exactly as a narrower distribution should be less uncertain. Hold the variance fixed and the Gaussian wins; change it and you are comparing different problems. This is also why entropy never appears without a variance constraint attached.
Where this shows up
Families under every model
Uncertainty in evaluation
Confidence intervals for a benchmark score are Student-t intervals whenever the sample is small and the variance is estimated from the same data. Heavy tails are why a handful of outlier examples can move a metric, and why robust losses and trimmed means matter in model evaluation.
Noise models in optimisation
Least-squares fitting assumes Gaussian residuals, which turns a maximum-likelihood fit into a sum of squares and makes the normal equations linear. When residuals are heavy-tailed, a Student-t or robust kernel is the honest model; that choice is exactly the loss-function trade discussed in nonlinear optimisation.
Gaussian belief states
A Kalman filter represents a robot's belief about its pose as a Gaussian, because a Gaussian is closed under the linear updates of prediction and correction. The exponential and gamma families appear on the other side of the same systems, as the waiting times between events in a Poisson arrival process.
Densities under change of variable
Pushing a uniform through a map is the standard way to sample any family, and the Jacobian that rescales the density is the bridge between the two views. That computation is developed in calculus of probability and used constantly in simulation.
Cheat sheet
Every formula in one place
| Family | Density | Mean / variance | When it appears |
|---|---|---|---|
| Uniform (a,b) | $\frac{1}{b-a}$ on [a,b] | $\frac{a+b}{2},\ \frac{(b-a)^2}{12}$ | Bounded range, total ignorance. |
| Exponential $(\lambda)$ | $\lambda e^{-\lambda x}$, $x\ge0$ | $\frac{1}{\lambda},\ \frac{1}{\lambda^{2}}$ | Memoryless waiting, Poisson arrivals. |
| Gamma $(k,\theta)$ | $\frac{x^{k-1}e^{-x/\theta}}{\Gamma(k)\theta^{k}}$ | $k\theta,\ k\theta^{2}$ | Time of the k-th arrival; sums of exponentials. |
| Beta $(\alpha,\beta)$ | $\frac{x^{\alpha-1}(1-x)^{\beta-1}}{B(\alpha,\beta)}$ | $\frac{\alpha}{\alpha+\beta},\ \frac{\alpha\beta}{(\alpha+\beta)^2(\alpha+\beta+1)}$ | Unknown proportion on [0,1]. |
| Student-t $(\nu)$ | $\frac{\Gamma(\frac{\nu+1}{2})}{\sqrt{\nu\pi}\,\Gamma(\frac{\nu}{2})}(1+\frac{x^2}{\nu})^{-\frac{\nu+1}{2}}$ | $0\ (\nu>1),\ \frac{\nu}{\nu-2}\ (\nu>2)$ | Small-sample mean; heavy tails. |
| Normal $(\mu,\sigma^2)$ | $\frac{1}{\sigma\sqrt{2\pi}}e^{-(x-\mu)^2/(2\sigma^2)}$ | $\mu,\ \sigma^{2}$ | CLT limit; maximum entropy at fixed variance. |
| Memorylessness | $P(X>s+t\mid X>s)=P(X>t)=e^{-\lambda t}$ | — | Only the exponential forgets how long you waited. |
| Gaussian entropy | $h=\tfrac12\log(2\pi e\,\sigma^{2})$ nats | — | The ceiling every fixed-variance density sits below. |
Further reading
Where to go deeper
- Larry Wasserman, All of Statistics, chapters 2 and 5 — a compact tour of the continuous families and the sampling distributions built from them.
- David MacKay, Information Theory, Inference, and Learning Algorithms, chapter 20 — maximum entropy as the principled way to choose a distribution from constraints.
- Thomas Cover and Joy Thomas, Elements of Information Theory, chapter 12 — the full derivation that the Gaussian maximises entropy at a fixed variance.
- E. T. Jaynes, Probability Theory: The Logic of Science, chapters 7 and 11 — the maximum-entropy argument and the Gaussian's special role.
- Grant Sanderson, "Why the Gaussian is so common", 3Blue1Brown — the central limit picture, drawn.