Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

The question

A small zoo of named distributions

Every discrete random variable is, in principle, just a table of probabilities: a number for each integer value. Nothing stops you from writing down an arbitrary table and calling it a distribution. In practice, though, the same few tables keep appearing, because the same few experiments keep appearing. A single yes-or-no trial, a fixed batch of independent trials, a run of trials stopped at a success, a count of rare events over time — these four stories cover a startling amount of applied probability. The named families are simply the probability tables those stories produce.

This part is about the zoo, but it is not a catalogue to memorise. The point of grouping the families is that they are related. The Bernoulli is the atom; the binomial is a sum of Bernoullis; the geometric is a binomial that stopped early; the negative binomial is a sum of geometrics; the Poisson is what the binomial becomes when the number of trials grows and the success probability shrinks in a coordinated way. Watching those relationships on a single canvas is worth more than any table, because it tells you when to reach for which family and why the formulas have the shape they do.

There is also a practical thread running through the whole act. Part 7 gave expectation and Part 8 gave variance; each family here carries its mean and variance as part of its identity, read straight off the panel as you move the sliders. The mean is where the mass sits; the variance measures how far it spreads; and the ratio between them is often the quickest way to recognise a family in the wild. A count whose variance equals its mean is a Poisson in disguise. A count that clusters tightly around its mean relative to that mean is binomial-like. Learning to read those signatures is a skill, not a formula.

By the end, the panel below should feel like a single object with six faces: Bernoulli, binomial, geometric, negative binomial, Poisson, and the limit that connects the last two. You will be able to switch between them, set each one's parameters, and watch the bars, the mean and the variance move together. Then we use the same panel for the limit animation: fix the product np, let n grow, and watch the binomial bars melt onto the Poisson bars.

💡 By the end of this part you'll see why the five discrete families are one family of stories in different light — how each arises, what its parameters mean, how the mean and variance encode its identity, and why the Poisson is the shadow the binomial casts when its trials become numerous and rare.
2

Bernoulli and binomial

One coin, then a batch of coins

The smallest interesting random variable takes just two values. A Bernoulli trial succeeds with probability p and fails with probability 1-p. Encode success as one and failure as zero, and the whole distribution is two numbers. Its mean is p, its variance is p(1-p), and both are already visible in the formula: the mean is the success probability itself, and the variance is largest at p=1/2, where the outcome is most uncertain.

$$P(X=k)=p^{k}(1-p)^{1-k},\qquad k\in\{0,1\},\qquad \mathbb{E}[X]=p,\qquad \operatorname{Var}(X)=p(1-p).$$

Repeat the trial n times, independently, and count the successes. The result is a binomial random variable. Each sequence of outcomes with exactly k successes and n-k failures has probability p^{k}(1-p)^{n-k}, and there are $\binom{n}{k}$ such sequences, so the probabilities multiply and add. The binomial is the first place a combinatorial coefficient appears inside a probability, which is exactly why Part 2 spent so long on counting.

$$P(X=k)=\binom{n}{k}p^{k}(1-p)^{n-k},\qquad k=0,1,\dots,n,\qquad \mathbb{E}[X]=np,\qquad \operatorname{Var}(X)=np(1-p).$$

The mean and variance tell a clean story. Every one of the n trials contributes a Bernoulli with mean p and variance p(1-p), and since the trials are independent the means and variances simply add — that is linearity of expectation from Part 7, and the variance addition rule from Part 8 doing its job. Nothing about the binomial's shape needs to be memorised: it is a sum, and the sum's first two moments are the sums of the parts.

Push the sliders and watch the shape. At p=1/2 the bars are symmetric, peaked at n/2. Move p toward zero and the mass piles up against the left wall, because successes are rare and the most likely count is small. Raise n with p fixed and the bars broaden and then, seen from a distance, start to look like a bell — the Central Limit Theorem of Part 18 is lurking there, but for now the useful fact is just that increasing n adds trials and therefore adds both mean and variance in proportion.

The distribution is named for Jacob Bernoulli, and it is the workhorse of every yes/no measurement: how many of n patients recover, how many of n packets arrive, how many of n bits flip. Whenever you can argue that the trials are independent and share the same success probability, the count of successes is binomial and its mean and variance are already known.

Bars are P(X=k) for the selected family. In limit mode the wide pale bars are $\text{Poisson}(\lambda)$ and the narrow solid bars are $\text{Binomial}(n,\lambda/n)$.

Switch the panel to Bernoulli and you will see just two bars; switch back to binomial and watch the same underlying coin reappear as a whole distribution of counts. That is the first relationship in the zoo: the binomial is what you get by summing Bernoulli trials, and the Bernoulli is what you get by summing exactly one.

3

Geometric and negative binomial

Stop at the first success, or the r-th

Change the experiment instead of the count. Flip until the first head and record how many flips it took. The result is the geometric distribution. For the run to last exactly k flips, the first k-1 must be failures and the k-th a success, so the probability is a geometric sequence in k.

$$P(X=k)=(1-p)^{k-1}p,\qquad k=1,2,3,\dots,\qquad \mathbb{E}[X]=\frac{1}{p},\qquad \operatorname{Var}(X)=\frac{1-p}{p^{2}}.$$

The mean 1/p is the one that surprises people: with a fair coin you wait two flips on average, and with a one-in-a-hundred event you wait a hundred. The variance is large and grows like 1/p^2, so rare successes are not merely delayed but erratic. The geometric distribution has the longest tail of the four families here, and that tail is where its famous property lives.

The memoryless property says that conditional on having already waited j failures, the remaining wait has the same distribution as a fresh start. In symbols, $P(X>j+k\mid X>j)=P(X>k)$. The coin does not remember the failures. This is the discrete twin of the exponential distribution's memorylessness, and Part 11 makes the link exact: the exponential is the continuous-time limit of the geometric, arrived at by taking the trial spacing to zero. If a success rate is constant in time and nothing ages, the wait is geometric in discrete steps and exponential in continuous time — the same idea wearing two clocks.

Now wait for the r-th success rather than the first. Summing r independent geometric waits gives the negative binomial. In the trials-until-r-successes convention, the last trial must be a success and the preceding k-1 contain r-1 successes in any order, which accounts for the binomial coefficient.

$$P(X=k)=\binom{k-1}{r-1}p^{r}(1-p)^{k-r},\qquad k=r,r+1,\dots,\qquad \mathbb{E}[X]=\frac{r}{p},\qquad \operatorname{Var}(X)=\frac{r(1-p)}{p^{2}}.$$

Setting r=1 recovers the geometric, exactly as it should. The mean multiplies by r, because the expected wait for r successes is r times the expected wait for one, but the variance also multiplies by r: the waits add and, by the variance rule, so do their variances. The panel holds the equivalent failure-count convention, in which the variable counts the failures before the r-th success; the two differ by a shift of r, which is why the panel's mean reads r(1-p)/p rather than r/p. The shape is the same, and the relationships are the same.

Negative binomials show up wherever you over-disperse a count. A Poisson count has variance equal to its mean; real counts of, say, accidents or purchases are often far more variable than that. The negative binomial has variance r(1-p)/p^2 which exceeds its mean r(1-p)/p whenever p<1, so it is the natural drop-in replacement when the Poisson is too tidy. Switch the panel between Poisson and negative binomial at matching means and compare the spreads: the negative binomial bars carry visibly more mass in the far tail.

4

Poisson

Counting rare events over a window

Imagine events arriving in time or space — radioactive decays, phone calls, typos, cosmic rays — with no memory and no clustering, and ask how many land in a fixed window. The Poisson distribution is the answer, with a single parameter $\lambda$, the expected count. Its probabilities come from a factorial in the denominator and powers of $\lambda$ in the numerator, and the normalisation is the Taylor series of the exponential.

$$P(X=k)=\frac{e^{-\lambda}\lambda^{k}}{k!},\qquad k=0,1,2,\dots,\qquad \mathbb{E}[X]=\lambda,\qquad \operatorname{Var}(X)=\lambda.$$

The signature of the Poisson is the equality of mean and variance. That single fact is enough to recognise it in data: compute the average count, compute the variance of the counts, and if the two agree you are looking at a memoryless arrival process. If the variance is much larger, reach for the negative binomial instead. The slider controls $\lambda$ directly, so the mean you read off the panel is also the variance, and the whole distribution shifts and spreads together.

Several operations preserve the family. Sums of independent Poissons are Poisson, with rates adding: two independent streams each Poisson merge into one Poisson whose rate is the sum. Thinning a Poisson — keeping each event independently with probability p — leaves a Poisson with rate $p\lambda$. Those two closure properties are why Poisson models compose so cleanly, and why they underpin queueing theory, photon counting and the arrival models in inference serving, where request arrivals and token emissions are routinely modelled as Poisson processes.

Notice what the Poisson does not require. It does not need a fixed number of trials; the count can in principle run to infinity, though the tail is thin and probabilities fall off like $\lambda^k/k!$. It does not need a unit of time; $\lambda$ is whatever brings the expected count to the observed window. And it does not need the events to be simultaneous, because a memoryless point process forbids two arrivals at the same instant. Later parts put these assumptions on firm footing through the Poisson process, but for now the panel shows the distribution itself: one knob, a mean equal to the variance, and a shape that starts skewed at small $\lambda$ and becomes symmetric as $\lambda$ grows.

The factorial in the denominator is doing real work. It is what makes the Poisson the limiting law for rare events, because it emerges from the binomial coefficient once the combinatorics are rescaled — and that rescaled limit is the next section.

5

The Poisson limit of the binomial

Many trials, tiny probability, fixed product

Here is the relationship that earns the Poisson its place. Take a binomial with n trials and success probability p, and hold the product $np=\lambda$ fixed while n grows and p shrinks. Each individual trial becomes rarer, but the expected number of successes stays put. In that regime the binomial pmf converges pointwise to the Poisson pmf.

$$\binom{n}{k}p^{k}(1-p)^{n-k}\;\longrightarrow\;\frac{e^{-\lambda}\lambda^{k}}{k!}\qquad\text{as }n\to\infty,\; p\to 0,\; np\to\lambda.$$

The proof is a two-line estimate with the algebra already in front of you. Write $p=\lambda/n$, expand the binomial coefficient as $n(n-1)\cdots(n-k+1)/k!$, and let n grow. The falling product divided by n^k tends to one, the factor $(1-\lambda/n)^{n}$ tends to $e^{-\lambda}$, and the leftover $(1-\lambda/n)^{-k}$ tends to one. What survives is $\lambda^k/k!$ times $e^{-\lambda}$. The binomial coefficient has become the 1/k! that defines the Poisson.

The panel's limit mode puts this on the canvas. The rate $\lambda$ is fixed by the slider; the animation raises n from a handful to a couple of hundred while setting $p=\lambda/n$ so the product never moves. The narrow solid bars are the binomial and the wide pale bars are the Poisson. At small n the two disagree: the binomial lives on a bounded range and its bars are chunky. As n climbs, the chunky bars split and migrate, the range widens, and the solid bars settle onto the pale ones. The readout reports the largest gap between the two pmfs, and you can watch it fall toward zero.

There is a second way to read the result, which is why it is sometimes called the law of rare events. Rare successes among many trials behave like independent arrivals in a window, and the only parameter that matters is their expected count. The n and the p that produced that count are individually irrelevant in the limit; only the product survives. That is exactly the collapse of two knobs into one that the panel shows, and it is why the Poisson needs no trial count at all.

The same rescaled-coefficient trick reappears throughout probability, and it connects back to this volume's vocabulary. The geometric's memoryless wait becomes the exponential under the same kind of scaling in Part 11, and the binomial's Gaussian shadow under a different scaling is the Central Limit Theorem in Part 18. One combinatorial family, three different limits depending on which knob goes to infinity. The Poisson limit is the rarest-events one: fix the expected count, send the trials to infinity, and the count of successes you cannot avoid meets you on the other side.

6

Where this shows up

The same five tables, under many names

Math

Counting capstones

Binomial probabilities are counting problems in disguise: the coefficient $\binom{n}{k}$ is the whole content of the counting part. Every time you choose which k of n trials succeed, you are doing the same enumeration that underlies combinations.

AI / ML

Token sampling and arrivals

A language model emitting tokens one at a time can be read as a sequence of Bernoulli choices, and the count of accepted tokens in speculative decoding is binomial. Request arrivals at a serving endpoint are modelled as a Poisson process, which is what sizes the batch in the decode loop.

Vision

Inlier counts in robust fitting

RANSAC reasons about the number of good correspondences in a random sample, a binomial count, and about the expected number of false matches over many hypotheses, which in the rare-error regime is Poisson. Both feed the iteration budget in multi-view geometry.

AI / ML

Likelihoods for counts

A Poisson or negative-binomial likelihood scores how surprising a count is, and it appears in count-based language modelling, event prediction and evaluation. The mean-equals-variance diagnostic is the first thing to check when choosing between them, as in language models.

7

Cheat sheet

Every formula in one place

FamilyPMFMean / varianceStory
Bernoulli$p^{k}(1-p)^{1-k},\;k\in\{0,1\}$$p\;;\;p(1-p)$One trial, success or failure.
Binomial$\binom{n}{k}p^{k}(1-p)^{n-k}$$np\;;\;np(1-p)$Successes in n independent trials.
Geometric$(1-p)^{k-1}p,\;k\ge1$$1/p\;;\;(1-p)/p^{2}$Trials until the first success; memoryless.
Negative binomial$\binom{k-1}{r-1}p^{r}(1-p)^{k-r},\;k\ge r$$r/p\;;\;r(1-p)/p^{2}$Trials until the r-th success; sum of r geometrics.
Poisson$e^{-\lambda}\lambda^{k}/k!$$\lambda\;;\;\lambda$Rare events in a window; mean equals variance.
Poisson limit$\binom{n}{k}p^{k}(1-p)^{n-k}\to e^{-\lambda}\lambda^{k}/k!$ when $np\to\lambda$$np\to\lambda\;;\;np(1-p)\to\lambda$Many trials, tiny probability, fixed expected count.
Memorylessness$P(X>j+k\mid X>j)=P(X>k)$—Geometric now; exponential in Part 11.
8

Further reading

Where to go deeper

9

Check your understanding

0/6 answered