Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

The question

What is in a list of numbers?

A single sample is conceptually simple and practically useless. It might be a hundred sensor readings, the token lengths of a day of requests, or the measured heights of a cohort. Before any inference, before any test, before any model, you need to answer a smaller question: what does this cloud of values look like? Human beings cannot read a thousand numbers, and a table of them hides exactly the features that matter — where the mass sits, how spread out it is, whether it is symmetric, whether a few extreme values are dragging everything else around.

So we compress. We draw a picture, and we quote a few summary numbers. Both steps are lossy, and being honest about what is lost is most of what separates a careful analysis from a careless one. A histogram replaces each value by the bin it lands in; a density estimate replaces it by a small smooth bump; a cumulative curve keeps the full order information but throws away the exact spacing. A mean reduces the whole sample to one number that every observation is allowed to influence. A median reduces it to one number that only the middle of the sorted list decides.

This part is about those choices, made visible. You will build a histogram whose bin width you control, overlay a kernel density estimate whose bandwidth you control, and read quantiles off a cumulative curve that has no tuning parameter at all. Then you will take one point and drag it into the tail, and watch the mean and median disagree about what "typical" means. Everything that follows in this series assumes you can look at a sample and see its shape; this is the part where you learn to.

💡 By the end of this part you'll see why a histogram and a KDE are estimates that depend on a smoothing knob, why the ECDF needs no knob and hands you the quantiles directly, and why one outlier can move the mean across the picture while the median barely notices.
2

The shape of one sample

A histogram, a kernel, and one knob each

The histogram is the oldest picture in statistics. Choose a set of equal-width bins covering the data, count how many values fall in each, and draw a bar of that height. It is a density estimate: the bar heights, divided by the total count and the bin width, estimate the probability density at that location. Everything about it is set by one decision, the bin width. Widen the bins and the picture gets smoother and more stable, but real features — a second hump, a hard edge, a gap — get smeared into their neighbours. Narrow the bins and the fine detail reappears, but so does the noise of a handful of points per bar, and the shape changes every time you shift where the first edge sits.

That is a bias–variance trade-off wearing work clothes. A wide bin averages many observations, so its estimate has low variance but is biased toward the local average, erasing structure. A narrow bin tracks structure but is based on almost no data, so its variance is enormous. There is no bin width that is both faithful and stable, because the estimator has a knob and the knob controls exactly that exchange. The same trade-off appears in every smoothing method in this series, and it is worth seeing it first in something as homely as a bar chart.

The kernel density estimate attacks the same problem differently. Instead of stacking values into boxes, place a small smooth bump — a kernel, usually a Gaussian — on top of each observation, and add the bumps together. The result is a continuous density rather than a staircase of bars:

$$\hat f_h(x) = \frac{1}{n h}\sum_{i=1}^{n} K\!\left(\frac{x - x_i}{h}\right)$$

Here $h$ is the bandwidth, the width of each bump, and it plays the role the bin width played before. Small $h$ gives a spiky curve that loyally follows every observation and its noise; large $h$ gives a smooth curve that may have smoothed away the interesting part. The choice is again a bias–variance trade-off, and again it is yours to make. A standard starting point is Silverman's rule of thumb, $h \approx 0.9\,\min(\hat\sigma,\, \mathrm{IQR}/1.349)\,n^{-1/5}$, which balances the two. It is a good default and nothing more; on skewed or multi-modal data it can visibly over- or under-smooth.

The demo below draws both pictures of the same sample. The bars and the curve are estimating the same thing and they should roughly agree, but move either slider and watch how the apparent shape — the number of humps, the sharpness of the shoulder on the right — changes without the data changing at all. The rug of ticks along the baseline is the ground truth: every data point is there, and the two smoothed estimates are just different ways of turning that discrete set into a continuous picture.

Bars: histogram density. Curve: kernel density estimate. Ticks: the actual observations. Both estimates depend on a smoothing knob you control.

⚠ A smoother picture is not a truer one. Widen the bins and the histogram looks calm because you have averaged away the noise, not because you have found the signal. The curve that best represents the sample is the one whose knob you can justify, and the honest thing to do is to look at several settings rather than one. This is the first appearance of a theme that runs through the whole series: every estimate in statistics is an estimate, and every estimate has a tuning choice lurking in it.
3

The ECDF and quantiles

The cumulative view, with no knob to turn

There is a third view that asks for no choice at all. Sort the sample, then define the empirical cumulative distribution function as the fraction of observations less than or equal to $x$:

$$F_n(x) = \frac{1}{n}\sum_{i=1}^{n} \mathbf{1}\{x_i \le x\}$$

It is a staircase that starts at zero, ends at one, and jumps by $1/n$ at each data point. Nothing is smoothed, nothing is binned, and no bandwidth has to be chosen. Every value the sample took is a step, in order, and the height at any point is the proportion of the sample at or below it. If you want to know what fraction of requests were under a latency budget, you read it off the height of the curve; if you want to know the value below which ninety-five percent of the data fall, you read off the side.

That second reading is the quantile, and it is the inverse of the ECDF. The median is the 0.5-quantile: the value with half the data on each side. The quartiles split the sample into quarters, and the interquartile range, $Q_3 - Q_1$, is a measure of spread that ignores both tails entirely. Percentiles are the same object at a finer resolution, and they are how latency and performance numbers are almost always quoted. All of them come for free from the ECDF, with no bin width or bandwidth standing between the data and the answer.

The ECDF also shows why the histogram and the density estimate agree in the end: the ECDF is exactly the running total of the histogram. Give each bar its height and stack the bars from left to right, and you have built the staircase. With coarse bins the cumulative histogram approximates the ECDF in visible steps; as the bins multiply it converges onto the same curve. The demo draws both at once for exactly that reason — the faint staircase is the cumulative histogram, the bold one is the ECDF, and the gap between them is the cost of binning.

Drag the percentile slider to move a horizontal line and read where it meets the curve. Below the line sits exactly that fraction of the data; the value on the horizontal axis is the corresponding quantile. The median is the crossing at one half; the 0.95 point is where tail behaviour starts to show, and it will matter again in Part 3 when the tails of one variable are compared with those of another.

Solid: the ECDF of the sample, a staircase with a step at every observation. Faint: the cumulative of a 12-bin histogram, which converges to it. The horizontal line marks a chosen percentile.

4

Mean, median, and one bad point

Which single number, and what it costs

Pictures are only half of a description; the other half is the summary number you quote in a sentence. The mean adds every value and divides by how many there are, so every observation gets a vote, and a single wild value gets a vote as large as any other. The median sorts the sample and takes the middle, so a value's rank matters and its magnitude does not; the largest number in the world, moved around, leaves the median unchanged as long as it stays largest. The mode is the most common value, or the peak of the density estimate, and is the natural summary for data that cluster rather than spread.

These three can disagree violently. The demo below is a small, ordinary-looking dataset with one point coloured differently. Drag that point to the right and watch the two vertical markers. The median moves slowly and then stops, because pushing the point further right does not change which value sits in the middle. The mean keeps sliding right, without limit, because it is dragged by the magnitude of the outlier. The gap between the two markers is a direct measure of how much the extreme value has distorted the average. Pull the point far enough and the mean can be larger than every honest observation in the sample.

This is why robustness is a word with a technical meaning. An estimator's breakdown point is the fraction of the data that must be corrupted before the estimate can be made arbitrarily wrong. For the mean it is zero: one bad observation is enough. For the median it is a half: you would have to corrupt half the sample before the median could be pushed anywhere. The IQR and the median absolute deviation share that robustness; the standard deviation and the mean do not. When a sensor glitches, a log line records a bogus giant value, or a measurement is simply mis-entered, robust summaries survive and non-robust ones do not.

Robustness is not free, and it is not always what you want. The mean uses all the information in the sample, so when the data really are clean it is more efficient than the median: it varies less from sample to sample. The choice is a judgement about your data, not a rule. If you believe the extreme values, use the mean; if you suspect them, use the median; and either way, say which one you used.

The shape of the sample explains the disagreement in advance. In a right-skewed sample, with a long tail of large values, the mean sits to the right of the median, and the distance between them grows with the skew; in a left-skewed sample it sits to the left. A box plot, the compact drawing built from the five-number summary — minimum, $Q_1$, median, $Q_3$, maximum, usually with whiskers and separately marked outliers — is a portrait of exactly this, and it is often paired with the mean drawn as a separate marker, precisely so the two can be compared at a glance.

Nine fixed values plus one point you can drag. The solid marker is the mean, the dashed one the median. Drag the outlier right and the mean runs away while the median holds its ground.

The base sample is fixed and symmetric; only the last point moves. The mean's breakdown point is zero, the median's is one half.

There is a last warning here, and it sets up the rest of this act. A handful of summary numbers is not the same thing as the data. Two samples can share a mean, a variance, a correlation, and a first regression line and still look nothing like each other — this is the point of Anscombe's four datasets, and it is why every statistician should plot before quoting. Summaries are useful precisely because they throw information away; the danger is forgetting which information, and trusting a number you have not looked at. Part 3 takes the next step and asks what happens when there are two variables, where an ellipse of covariance is the new picture and correlation is the new summary that hides more than it shows.

5

Where this shows up

The same summaries, two very different worlds

ML / AI

Latency tails and percentiles

Serving systems report p50, p95 and p99 latency rather than a mean, because request latency is heavy-tailed and the few slowest requests dominate the average. A mean latency of 40 ms can sit on top of a p99 of two seconds, and it is the p99 the users actually feel. The LLM serving metrics chapter treats this directly: the ECDF and its quantiles are the honest summary of a tail, and the mean is the number that hides it.

Robotics

Robust summaries of noisy measurements

A robot estimating its displacement from wheel encoders or visual odometry accumulates occasional gross errors: a slip, a mismatched feature, a dropped frame. Averaging those readings lets one blunder corrupt the whole estimate, which is why pipelines reach for medians, trimmed means, and RANSAC-style consensus instead. That robust side of the mean-versus-median question is what you have just dragged with the mouse.

The same fork appears everywhere data are summarised. A monitoring dashboard that reports only the average of a heavy-tailed quantity is a histogram with one bar. A benchmark score is a mean over a sample of inputs, and its stability is a question about that sample. A model's loss curve is a sequence of summaries, each one throwing away the distribution of per-example errors. The method you choose to compress a sample into a sentence is never neutral; it decides which parts of the world you can see.

Further reading

The references below all treat description as estimation, which is the attitude this series takes. If you take away one thing, take away the picture of the same data under one knob — bin width, bandwidth, or quantile — and the fact that the ground truth underneath it never moves.

Silverman's monograph is the canonical treatment of density estimation and the source of the bandwidth rule of thumb used in the demo; Wasserman covers histograms, the ECDF and quantiles compactly in the opening chapters. Anscombe's four datasets are worth looking at once and remembering forever, and Hunter's short paper makes the case that a histogram is itself an estimate with a bias–variance trade-off.

Cheat sheet

TermMeaning here
HistogramDensity estimate from counts in fixed-width bins; its one choice is the bin width
Bin widthWide bins smooth and bias the shape; narrow bins track detail and add variance
KDESum of small kernels centred on the data; its one choice is the bandwidth $h$
Bandwidth $h$Width of each kernel bump; small is spiky, large is oversmooth
Silverman's rule$0.9\min(\hat\sigma, \mathrm{IQR}/1.349)n^{-1/5}$, a reasonable default for $h$
ECDF $F_n$Fraction of the sample at or below $x$; a staircase, with no tuning parameter
QuantileInvert the ECDF: the $p$-quantile is the value with fraction $p$ of the data below it
MedianThe 0.5-quantile; decided by rank, so outliers do not move it
MeanSum divided by count; every value votes, so one outlier can move it without bound
IQR$Q_3 - Q_1$; robust spread, and the width of the box in a box plot
SkewRight-skewed data put the mean above the median; left-skewed data put it below
6

Check your understanding

0/4 answered