From mass to density
A discrete random variable hands you a number for every outcome: the chance of rolling a three is one sixth, full stop. A continuous random variable cannot do that. The chance that a bus arrives at exactly 8:03:17.421 while you wait is not small — it is zero, and the same is true of every individual instant. Yet waiting times clearly have a shape: short waits are common and long waits are rare. Probability has not disappeared; it has spread itself out over the line, and the thing that carries it is the density. This part builds the density as the limit of a histogram, then insists on the one fact that trips everyone up: a density is not a probability, and a density can be larger than one. What is always a probability is the area under it.
The question
Why nothing has probability zero but still happens
Suppose you sample a number from a continuous distribution — the exact position of a dart on a board, the exact time until a radioactive atom decays, the exact weight of a manufactured bolt. Ask for the probability that the result equals some pre-chosen value x. The honest answer is zero. There are uncountably many possible values, no single one carries any weight by itself, and no amount of rescaling changes that. If point probabilities were all we had, the continuous world would look degenerate.
The resolution is to stop asking about points and start asking about intervals. Even though P(X=x)=0, the chance that X lands between 8:03 and 8:04 is a perfectly good positive number, and so is the chance that a bolt weighs between 9.98 and 10.02 grams. Intervals have width, and width is what a continuous distribution can hand out. The density is the gadget that tells you how much probability sits per unit of the line at each point, and then an integral accumulates that rate over an interval to produce an actual probability.
This is exactly how mass density works in physics, and the analogy is worth taking literally. A steel rod has a mass density at every point, in grams per centimetre. Ask for the mass at a point and the answer is zero, because a point has no length; the density is not a mass, it is a rate. Ask for the mass of the segment between two marks and you integrate the density over the segment. Probability density is the same idea with probability as the mass: f(x) is probability per unit length, and only $\int_a^b f(x)\,dx$ is a probability.
So the plan for this part. First we take a finite sample and draw its histogram, then let the bins narrow until the jagged skyline settles onto a smooth curve — that curve is the density, and it is what your histogram was approximating all along. Then we confront the consequence nobody expects: because the density is a rate, it can take values far above one, and a narrow spike does exactly that. Finally we shade an interval and watch its area match the difference of the cdf at the two endpoints, which is the operational meaning of the whole construction.
From histogram to density
Narrow the bins and the skyline becomes a curve
Here is a way to see a density without any calculus at all. Draw ten thousand samples from a continuous distribution you can sample but cannot write down, and sort them into bins of width h. Above each bin draw a rectangle whose area equals the fraction of samples that fell inside it. That is a histogram with the counts normalised by n and by the bin width, and the area of each rectangle is then an empirical probability of landing in that bin.
The total area is always one, no matter how wide or narrow the bins are, because the fractions always sum to one and we divided by the bin width exactly once. Now shrink h. The rectangles get thinner and taller, the staircase gets finer, and if the underlying variable has a smooth distribution the skyline converges to a fixed curve. That limit curve f is the probability density function, and the fact that its integral is one is inherited directly from the fact that the histogram areas always summed to one.
Play with the slider. Coarse bins give a blocky approximation you can read left to right; fine bins give a noisy but faithful outline. Push the bins too fine and the histogram stops improving — with a fixed sample it starts to jitter, because each thin bin holds too few points to give a stable height. The density is a limit in two directions at once: infinitely many bins and infinitely many samples. With finite data you get to pick a width that trades bias against noise, which is precisely the bias–variance trade every statistical estimator faces.
Flip the toggle to see counts instead of density. The shape is identical, but now the vertical axis measures raw frequencies and the total area is n, not one. Switching back to density rescales the rectangles so their areas are probabilities again. That rescaling by 1/(n h) is the entire difference between a picture of your data and an estimate of the underlying law.
Bars show a seeded sample of 6000 draws from N(0,1), normalised so each bar's area is a probability. The smooth pink curve is the true density $\varphi(x)$; narrow the bins and watch the bars converge onto it.
One more reading of the same picture. The density at x is the limit of the fraction of samples falling in a tiny window around x, divided by the width of that window. It is a rate: probability per unit length. This is why the units of a density are inverse units of X — a time density is measured in per-second, a weight density in per-gram. If you change the unit of measurement, the density numbers change, but the integrals do not.
The density is not a probability
Why f(x) is allowed to be bigger than one
Every probability lies between zero and one, so a first glance at a density that reaches 3.3 looks like a bug. It is not. The value of f at a point is not the probability of that point; it is the rate at which probability accumulates there. To compare it with a probability you must multiply by a length. The quantity $f(x)\,h$ is approximately the probability of landing in a window of width h around x, and that product is always below one for small enough h, however tall the density is.
The cleanest way to make peace with this is to squeeze the distribution. Take a beta distribution concentrated around one half — a bell of total area one, but narrow. Conservation of area forces the peak up as the width comes down. Halve the width and you roughly double the peak. A spike narrow enough has a peak as large as you like, while every actual probability it describes stays between zero and one. Nothing is violated; the peak is simply inheriting the units of inverse length.
Drag the concentration slider and watch the peak of a symmetric beta density. At low concentration the curve is broad and shallow; as concentration grows the same unit of area piles up and the maximum crosses one, then two, then higher. The dashed line at height one is the false ceiling that intuition wants to impose. Notice how the shaded interval's probability keeps changing in the ordinary way even as the peak soars — the probability depends on the area, not on the height.
There is a discrete cousin of this confusion worth naming. For a discrete variable the pmf p(k) is a probability, so it genuinely cannot exceed one. The density replaces a sum with an integral, and in doing so changes the object from an amount into a rate. The honest reading of a density axis is "probability per unit", and the honest question is never "how big is f here" but "how much area does f enclose there".
A symmetric $\text{Beta}(k,k)$ density on [0,1]. The dashed line is height one; the pink band is $P(a\le X\le b)=\int_a^b f$. Only the integral is a probability.
Keep the units in mind and the trap closes. If X is a time in seconds, a density of 3 per second means that in the next millisecond the chance of landing in a window is about 0.003 — an ordinary probability well below one. Change the units to milliseconds and the same curve's density value drops by a factor of a thousand; the probabilities it integrates to are untouched. A density is a bookkeeping device tied to a choice of ruler, while probability itself is not.
Area is probability
Two ways to compute the same number
Define the cumulative distribution function $F(x)=P(X\le x)$, the probability that the variable has landed no further right than x. Because it is built from intervals, it is the running integral of the density from the far left up to x. The fundamental theorem of calculus then runs the relation backwards: differentiating the accumulated area recovers the height of the curve. Density and cdf are two views of one object, related by integration in one direction and differentiation in the other.
The middle identity is the one to carry around. The probability of an interval is the area under the density between its endpoints, and it is equally the rise of the cdf across them. The first is a geometry fact and the second is an arithmetic one, and they agree because the cdf was defined as the accumulated area. Below, drag either edge of the shaded band. The shaded area on the upper curve and the height difference on the cdf both update together, and the readout confirms they are the same number.
Watch the endpoints move toward each other. As the band narrows, the cdf difference shrinks toward zero, which is the geometric statement that a single point carries no probability. Move the whole band to the tail and both the area and the rise become small, because the density is low out there. Put the band where the density peaks and the cdf climbs fastest, because that is where probability per unit length is highest. Everything you know about the shape of the density is a statement about the slope of the cdf.
Drag the two handles to move a and b. The shaded area is $P(a\le X\le b)$ and the readout shows it equal to F(b)-F(a).
This is why so many continuous calculations reduce to a cdf difference. The probability of any interval whatsoever — a tail, a central band, the union of two disjoint stretches — is a difference of F values at the right endpoints. Summing and integrating are the same operation wearing different clothes, and once the density is in hand the cdf is just its running total. The next part tabulates these integrals for the standard families, so that in practice you look them up rather than compute them.
Where this shows up
Densities under everything continuous
Integrals over regions
A probability is an integral of a density, and in several variables it becomes an integral over a region. That machinery — changing the order of integration, slicing a volume into areas — is developed in multiple integrals, and it is exactly what turns a joint density into a marginal one.
The calculus of probability
Push a continuous variable through a smooth map and the density rescales by a Jacobian. The geometric picture, with densities as areas preserved under a change of coordinates, is told in Calculus of probability; this part is the density it presumes.
The discrete contrast
Going the other way, a pmf assigns a mass to each value and the total is a sum. The families where that sum is easy — binomial, Poisson, geometric — are collected in discrete families, and comparing them with a density sharpens the difference between a mass and a rate.
Likelihoods and log-densities
Language models are trained by maximising a log density over a continuum of predictions, and per-token losses are log-densities that may be strongly negative even where probabilities are moderate. The connection between density, likelihood and loss is drawn out in language models.
Cheat sheet
Every formula in one place
| Idea | Formula | Reading |
|---|---|---|
| Density | $f(x)\ge 0$ | Probability per unit length of x; not a probability. |
| Normalisation | $\int_{-\infty}^{\infty} f(x)\,dx=1$ | Total area is one, inherited from the histogram. |
| Interval probability | $P(a\le X\le b)=\int_a^b f(x)\,dx$ | The shaded area between the endpoints. |
| Point probability | P(X=x)=0 | A point has zero width, so zero area. |
| Density above one | f(x)>1 is allowed | Narrow spikes have tall peaks; only integrals are probabilities. |
| cdf | $F(x)=\int_{-\infty}^{x} f(t)\,dt$ | The running total of area up to x. |
| Recovering the density | f(x)=F'(x) | Differentiate the cdf; density is its slope. |
| cdf difference | $P(a\le X\le b)=F(b)-F(a)$ | Area and rise are the same number. |
| Histogram limit | $\frac{\#\{x_i\in\text{bin}\}}{n\,h}\to f(x)$ | Area-normalised bars converge as bins narrow. |
Further reading
Where to go deeper
- David MacKay, Information Theory, Inference, and Learning Algorithms, chapter 2 — densities derived from the histogram limit, with the same units argument.
- William Feller, An Introduction to Probability Theory and Its Applications, volume II — the careful treatment of densities, support and absolute continuity.
- Sheldon Ross, A First Course in Probability, chapters 5–6 — continuous random variables, densities and the uniform, exponential and normal families.
- Larry Wasserman, All of Statistics, chapter 2 — density estimation, bin width and the bias–variance trade the histogram exhibits.
- Grant Sanderson, "The birthday paradox" and the probability playlist, 3Blue1Brown — visual instincts for continuous probability.