Describing one sample
A sample arrives as a flat list of numbers, and the first job is to make it legible. Three pictures have become standard: the histogram, the kernel density estimate, and the empirical cumulative distribution function. They are three views of the same data, and two of them force a choice on you — how wide to make the bins, how wide to make the kernel — that changes what you see without changing a single data point. The third choice is quieter: which single number you quote as the summary. As one unlucky value wanders off to the right, the mean chases it and the median does not, and the gap between them is the whole argument for robust statistics.
The question
What is in a list of numbers?
A single sample is conceptually simple and practically useless. It might be a hundred sensor readings, the token lengths of a day of requests, or the measured heights of a cohort. Before any inference, before any test, before any model, you need to answer a smaller question: what does this cloud of values look like? Human beings cannot read a thousand numbers, and a table of them hides exactly the features that matter — where the mass sits, how spread out it is, whether it is symmetric, whether a few extreme values are dragging everything else around.
So we compress. We draw a picture, and we quote a few summary numbers. Both steps are lossy, and being honest about what is lost is most of what separates a careful analysis from a careless one. A histogram replaces each value by the bin it lands in; a density estimate replaces it by a small smooth bump; a cumulative curve keeps the full order information but throws away the exact spacing. A mean reduces the whole sample to one number that every observation is allowed to influence. A median reduces it to one number that only the middle of the sorted list decides.
This part is about those choices, made visible. You will build a histogram whose bin width you control, overlay a kernel density estimate whose bandwidth you control, and read quantiles off a cumulative curve that has no tuning parameter at all. Then you will take one point and drag it into the tail, and watch the mean and median disagree about what "typical" means. Everything that follows in this series assumes you can look at a sample and see its shape; this is the part where you learn to.
The shape of one sample
A histogram, a kernel, and one knob each
The histogram is the oldest picture in statistics. Choose a set of equal-width bins covering the data, count how many values fall in each, and draw a bar of that height. It is a density estimate: the bar heights, divided by the total count and the bin width, estimate the probability density at that location. Everything about it is set by one decision, the bin width. Widen the bins and the picture gets smoother and more stable, but real features — a second hump, a hard edge, a gap — get smeared into their neighbours. Narrow the bins and the fine detail reappears, but so does the noise of a handful of points per bar, and the shape changes every time you shift where the first edge sits.
That is a bias–variance trade-off wearing work clothes. A wide bin averages many observations, so its estimate has low variance but is biased toward the local average, erasing structure. A narrow bin tracks structure but is based on almost no data, so its variance is enormous. There is no bin width that is both faithful and stable, because the estimator has a knob and the knob controls exactly that exchange. The same trade-off appears in every smoothing method in this series, and it is worth seeing it first in something as homely as a bar chart.
The kernel density estimate attacks the same problem differently. Instead of stacking values into boxes, place a small smooth bump — a kernel, usually a Gaussian — on top of each observation, and add the bumps together. The result is a continuous density rather than a staircase of bars:
Here $h$ is the bandwidth, the width of each bump, and it plays the role the bin width played before. Small $h$ gives a spiky curve that loyally follows every observation and its noise; large $h$ gives a smooth curve that may have smoothed away the interesting part. The choice is again a bias–variance trade-off, and again it is yours to make. A standard starting point is Silverman's rule of thumb, $h \approx 0.9\,\min(\hat\sigma,\, \mathrm{IQR}/1.349)\,n^{-1/5}$, which balances the two. It is a good default and nothing more; on skewed or multi-modal data it can visibly over- or under-smooth.
The demo below draws both pictures of the same sample. The bars and the curve are estimating the same thing and they should roughly agree, but move either slider and watch how the apparent shape — the number of humps, the sharpness of the shoulder on the right — changes without the data changing at all. The rug of ticks along the baseline is the ground truth: every data point is there, and the two smoothed estimates are just different ways of turning that discrete set into a continuous picture.
Bars: histogram density. Curve: kernel density estimate. Ticks: the actual observations. Both estimates depend on a smoothing knob you control.
The ECDF and quantiles
The cumulative view, with no knob to turn
There is a third view that asks for no choice at all. Sort the sample, then define the empirical cumulative distribution function as the fraction of observations less than or equal to $x$:
It is a staircase that starts at zero, ends at one, and jumps by $1/n$ at each data point. Nothing is smoothed, nothing is binned, and no bandwidth has to be chosen. Every value the sample took is a step, in order, and the height at any point is the proportion of the sample at or below it. If you want to know what fraction of requests were under a latency budget, you read it off the height of the curve; if you want to know the value below which ninety-five percent of the data fall, you read off the side.
That second reading is the quantile, and it is the inverse of the ECDF. The median is the 0.5-quantile: the value with half the data on each side. The quartiles split the sample into quarters, and the interquartile range, $Q_3 - Q_1$, is a measure of spread that ignores both tails entirely. Percentiles are the same object at a finer resolution, and they are how latency and performance numbers are almost always quoted. All of them come for free from the ECDF, with no bin width or bandwidth standing between the data and the answer.
The ECDF also shows why the histogram and the density estimate agree in the end: the ECDF is exactly the running total of the histogram. Give each bar its height and stack the bars from left to right, and you have built the staircase. With coarse bins the cumulative histogram approximates the ECDF in visible steps; as the bins multiply it converges onto the same curve. The demo draws both at once for exactly that reason — the faint staircase is the cumulative histogram, the bold one is the ECDF, and the gap between them is the cost of binning.
Drag the percentile slider to move a horizontal line and read where it meets the curve. Below the line sits exactly that fraction of the data; the value on the horizontal axis is the corresponding quantile. The median is the crossing at one half; the 0.95 point is where tail behaviour starts to show, and it will matter again in Part 3 when the tails of one variable are compared with those of another.
Solid: the ECDF of the sample, a staircase with a step at every observation. Faint: the cumulative of a 12-bin histogram, which converges to it. The horizontal line marks a chosen percentile.
Mean, median, and one bad point
Which single number, and what it costs
Pictures are only half of a description; the other half is the summary number you quote in a sentence. The mean adds every value and divides by how many there are, so every observation gets a vote, and a single wild value gets a vote as large as any other. The median sorts the sample and takes the middle, so a value's rank matters and its magnitude does not; the largest number in the world, moved around, leaves the median unchanged as long as it stays largest. The mode is the most common value, or the peak of the density estimate, and is the natural summary for data that cluster rather than spread.
These three can disagree violently. The demo below is a small, ordinary-looking dataset with one point coloured differently. Drag that point to the right and watch the two vertical markers. The median moves slowly and then stops, because pushing the point further right does not change which value sits in the middle. The mean keeps sliding right, without limit, because it is dragged by the magnitude of the outlier. The gap between the two markers is a direct measure of how much the extreme value has distorted the average. Pull the point far enough and the mean can be larger than every honest observation in the sample.
This is why robustness is a word with a technical meaning. An estimator's breakdown point is the fraction of the data that must be corrupted before the estimate can be made arbitrarily wrong. For the mean it is zero: one bad observation is enough. For the median it is a half: you would have to corrupt half the sample before the median could be pushed anywhere. The IQR and the median absolute deviation share that robustness; the standard deviation and the mean do not. When a sensor glitches, a log line records a bogus giant value, or a measurement is simply mis-entered, robust summaries survive and non-robust ones do not.
Robustness is not free, and it is not always what you want. The mean uses all the information in the sample, so when the data really are clean it is more efficient than the median: it varies less from sample to sample. The choice is a judgement about your data, not a rule. If you believe the extreme values, use the mean; if you suspect them, use the median; and either way, say which one you used.
The shape of the sample explains the disagreement in advance. In a right-skewed sample, with a long tail of large values, the mean sits to the right of the median, and the distance between them grows with the skew; in a left-skewed sample it sits to the left. A box plot, the compact drawing built from the five-number summary — minimum, $Q_1$, median, $Q_3$, maximum, usually with whiskers and separately marked outliers — is a portrait of exactly this, and it is often paired with the mean drawn as a separate marker, precisely so the two can be compared at a glance.
Nine fixed values plus one point you can drag. The solid marker is the mean, the dashed one the median. Drag the outlier right and the mean runs away while the median holds its ground.
There is a last warning here, and it sets up the rest of this act. A handful of summary numbers is not the same thing as the data. Two samples can share a mean, a variance, a correlation, and a first regression line and still look nothing like each other — this is the point of Anscombe's four datasets, and it is why every statistician should plot before quoting. Summaries are useful precisely because they throw information away; the danger is forgetting which information, and trusting a number you have not looked at. Part 3 takes the next step and asks what happens when there are two variables, where an ellipse of covariance is the new picture and correlation is the new summary that hides more than it shows.
Where this shows up
The same summaries, two very different worlds
Latency tails and percentiles
Serving systems report p50, p95 and p99 latency rather than a mean, because request latency is heavy-tailed and the few slowest requests dominate the average. A mean latency of 40 ms can sit on top of a p99 of two seconds, and it is the p99 the users actually feel. The LLM serving metrics chapter treats this directly: the ECDF and its quantiles are the honest summary of a tail, and the mean is the number that hides it.
Robust summaries of noisy measurements
A robot estimating its displacement from wheel encoders or visual odometry accumulates occasional gross errors: a slip, a mismatched feature, a dropped frame. Averaging those readings lets one blunder corrupt the whole estimate, which is why pipelines reach for medians, trimmed means, and RANSAC-style consensus instead. That robust side of the mean-versus-median question is what you have just dragged with the mouse.
The same fork appears everywhere data are summarised. A monitoring dashboard that reports only the average of a heavy-tailed quantity is a histogram with one bar. A benchmark score is a mean over a sample of inputs, and its stability is a question about that sample. A model's loss curve is a sequence of summaries, each one throwing away the distribution of per-example errors. The method you choose to compress a sample into a sentence is never neutral; it decides which parts of the world you can see.
Further reading
The references below all treat description as estimation, which is the attitude this series takes. If you take away one thing, take away the picture of the same data under one knob — bin width, bandwidth, or quantile — and the fact that the ground truth underneath it never moves.
Silverman's monograph is the canonical treatment of density estimation and the source of the bandwidth rule of thumb used in the demo; Wasserman covers histograms, the ECDF and quantiles compactly in the opening chapters. Anscombe's four datasets are worth looking at once and remembering forever, and Hunter's short paper makes the case that a histogram is itself an estimate with a bias–variance trade-off.
- Larry Wasserman, All of Statistics: A Concise Course in Statistical Inference, chapters 1–2 — the ECDF, quantiles and the empirical distribution.
- Bernard W. Silverman, Density Estimation for Statistics and Data Analysis — kernels, bandwidths and the rule of thumb.
- Francis Anscombe, "Graphs in Statistical Analysis", 1973 — four datasets with identical summaries and different shapes.
- J. Stewart Hunter, "The Exponential Weighted Moving Average" and the histogram-as-estimator literature — why bin width is a bias–variance choice.
- 3Blue1Brown, Essence of Statistics series — the visual style this guide follows.
Cheat sheet
| Term | Meaning here |
|---|---|
| Histogram | Density estimate from counts in fixed-width bins; its one choice is the bin width |
| Bin width | Wide bins smooth and bias the shape; narrow bins track detail and add variance |
| KDE | Sum of small kernels centred on the data; its one choice is the bandwidth $h$ |
| Bandwidth $h$ | Width of each kernel bump; small is spiky, large is oversmooth |
| Silverman's rule | $0.9\min(\hat\sigma, \mathrm{IQR}/1.349)n^{-1/5}$, a reasonable default for $h$ |
| ECDF $F_n$ | Fraction of the sample at or below $x$; a staircase, with no tuning parameter |
| Quantile | Invert the ECDF: the $p$-quantile is the value with fraction $p$ of the data below it |
| Median | The 0.5-quantile; decided by rank, so outliers do not move it |
| Mean | Sum divided by count; every value votes, so one outlier can move it without bound |
| IQR | $Q_3 - Q_1$; robust spread, and the width of the box in a box plot |
| Skew | Right-skewed data put the mean above the median; left-skewed data put it below |