t, z, chi-square, F and the likelihood-ratio
By now you have met a small crowd of tests, each with its own statistic, its own table of critical values, and its own letter. They are the same machine in different costumes. Every one of them reduces the data to a number measuring departure from the null, asks how large that number typically gets when the null is true, and reports the probability of a departure at least that large. Only the middle piece changes — the null distribution. A question about a mean runs into Student's t or the normal curve; a question about counts runs into chi-square; a question about whether a model earns its extra parameters runs into F through the likelihood-ratio. This part is a reference card for that zoo: pick a question and see the matching null distribution with the observed statistic marked and the p-value tail shaded.
The question
What do all these tests have in common?
A hypothesis test compares what you saw with what you would have seen if nothing interesting were going on. The comparison has three ingredients. First, a test statistic: a single number computed from the data, arranged to be near zero when the null story is true and large when the alternative is true. Second, a null distribution: the distribution that statistic would follow if the null were exactly true and the experiment repeated forever. Third, a tail: the region of that distribution corresponding to values at least as extreme as the one observed, whose probability is the p-value.
Only the first ingredient depends on your data. The third is just an integral over the second. That is why the null distribution is the real subject of this page: choose the right one and the test nearly writes itself, choose the wrong one and every number downstream is wrong in a way no amount of data will fix.
The four null distributions are not arbitrary. The normal, t, chi-square and F are all built from averages and sums of squares of normal variables, which is why the same four shapes keep reappearing across different subjects.
One statistic, many null distributions
Why the same recipe gives four different curves
Start with the most familiar test. To ask whether a mean differs from a value, standardise the difference by its own standard error:
If the observations are independent normal draws with unknown variance, this quantity follows a Student's t distribution with $n-1$ degrees of freedom. With a known variance the same construction would give the normal; the t is the normal with the uncertainty of the estimated spread folded in, which is why it has heavier tails that thin toward the normal as $n$ grows.
Change the question and the statistic changes shape. To ask whether observed counts match expected counts, sum the standardised squared discrepancies:
This is a sum of squared standardised pieces, and sums of squares are exactly what the chi-square distribution describes. Its degrees of freedom count the independent pieces of information, so a two-by-two table has one and a two-by-three table has two.
To ask whether one variance is larger than another, or whether group means vary more than noise would explain, take a ratio of two independent variance estimates:
The F distribution is the ratio of two independent chi-square variables, each divided by its degrees of freedom. It is always positive and right-skewed: a ratio of scales cannot be negative, and a few draws land far above one when the numerator is genuinely inflated.
So the four shapes are not four mysteries. They are the normal, and three things you can build from it: a ratio of a normal to an estimated spread (t), a sum of squares (chi-square), and a ratio of two such sums (F). Every test below uses one of these as its null distribution, each carrying its own degrees of freedom — the only parameter the shape needs.
The reference: pick a question
The signature interaction
The panel below is a reference card you can operate. Each button is a question a researcher might ask. Press one and the page simulates a dataset in the shape that question calls for, computes the appropriate statistic, draws the matching null distribution, marks the observed statistic, and shades the p-value tail. The table underneath prints the row a textbook would: which test, which statistic, which degrees of freedom, which p-value, and which assumptions the answer leans on.
The sliders change the sample size and the effect size and recompute everything. The effect slider controls how far the simulated truth is from the null, so sliding it up makes the statistic larger and the shaded tail smaller. The sample-size slider controls how tightly the statistic concentrates, producing the same qualitative effect by a different route. Together they are the two levers behind every power calculation in the series.
The null distribution for the chosen question. The heavy vertical line is the observed statistic; the shaded region is the p-value tail.
| Test | Statistic | Null distribution · df | p-value | Assumptions |
|---|---|---|---|---|
| Pick a question above to fill in the row. | ||||
The assumptions behind the shapes
What each curve quietly takes for granted
The null distribution is a claim about the world, and every claim has conditions. The t test assumes independent observations and a standardised mean that behaves like a normal over an estimated spread. At small samples that needs roughly normal data; as n grows, the central limit theorem rescues the numerator and the denominator settles down, so the test tolerates all but the most extreme skew and outliers. The z test assumes a known variance, rarely true in practice; it survives as the large-sample limit of the t.
The two-sample t test adds a choice. The pooled version assumes both groups share one variance, buying a clean degrees-of-freedom count. When that is doubtful, Welch's version estimates each variance separately and derives an approximate df from the sample sizes and spreads. Welch is the safer default and is the version used above; the cost is a fractionally less powerful test when equal variances really do hold.
Independence is the assumption most often violated and least often checked. If observations come in clusters — repeated measurements on one subject, several tokens from one document, frames from one sequence — the effective sample size is smaller than the count suggests and every standard error is too small. That failure cannot be fixed by picking a different table; it has to be designed away.
The chi-square test of independence assumes the entries are counts of independent events and that each expected cell count is not too small — the rule of thumb is at least five, so the discrete counting is well approximated by a continuous curve. The F test assumes independent, normally distributed errors and equal variance across groups. When equal variance fails there is a Welch-style analogue, and when the data are badly non-normal the honest move is often a permutation test, which rebuilds the null by shuffling labels rather than trusting a shape.
Notice the pattern in these caveats. None is about the test statistic; all are about whether the null distribution is the right curve to integrate. That is why this page insists on the three-part decomposition: it tells you where to look when a test feels like it might be lying.
The likelihood-ratio engine
One idea that produces all the others
The tests above can look like unrelated folklore. Underneath them is a single construction. Fit two models to the same data: a smaller null model, in which some parameter is fixed or held at a boundary, and a larger alternative model in which that parameter is free. Each model has a maximised log-likelihood, a number saying how plausible the data are under the best version of that model. Twice the difference is the likelihood-ratio statistic:
Wilks' theorem says that, under the null and some regularity conditions, this statistic is approximately chi-square distributed with degrees of freedom equal to the number of extra free parameters in the alternative model. That is the general engine: it converts any nested pair of models into a p-value with one command.
The familiar tests are what the engine produces in special cases. A one-sample t test compares a model with a free mean against one with the mean pinned to zero, and the likelihood-ratio statistic is a monotone function of $t^2$. The chi-square goodness-of-fit test is the engine applied to multinomial counts. The F test for nested regressions is a small-sample correction to the same idea. The z test is what remains when the sample is large enough that the correction disappears. None is more fundamental than the others; the likelihood-ratio makes the shared skeleton visible.
This is also why the language of degrees of freedom is everywhere. Each extra parameter the alternative spends costs one degree of freedom, and the statistic rewards a genuinely better fit while punishing one that improved only because more parameters were available. The probability guide is where the algebra behind Wilks' theorem lives; the intuition is simply that a better fit is only interesting if it is bigger than the free parameters can explain by chance.
When the regularity conditions fail — tiny samples, parameters on a boundary, heavily discrete outcomes — the asymptotic chi-square can be wrong, and the safe answer is to build the null distribution by simulation instead. The bootstrap and the permutation test in Part 12 are exactly that move: rebuild the distribution of the statistic under the null from the data at hand rather than trust a curve.
Where this shows up
The same three ingredients, two worlds
The F test in regression
In the least-squares part, adding predictors always lowers the residual sum of squares, even if the predictors are noise. The F test asks whether the drop is larger than the extra degrees of freedom can explain, which is the likelihood-ratio engine applied to nested linear models. Every regression printout's "F-statistic" line is this page's F distribution with two degrees-of-freedom counts in the denominator.
Which test for a benchmark
The evaluation chapter compares models on a held-out set, and the right test depends on the question: paired differences across items want a t test on the differences, success rates want a proportion test, and a comparison across several models at once wants an F or a likelihood-ratio. Choosing the statistic before choosing the null distribution is where most benchmark comparisons go quietly wrong.
The same three ingredients recur wherever a decision is made from noisy data. A robot asking whether its odometry has drifted is comparing a statistic against a null distribution of residuals; a serving team comparing two latency measurements needs the two-sample machinery rather than a bare difference of averages. Once the pattern is visible, "which test do I use?" becomes the concrete question "what is my statistic, and what does it do when the null is true?"
Further reading
The references below treat the tests as one family rather than a catalogue to memorise. If you take away one thing, take away the three-part decomposition — statistic, null distribution, tail — and the habit of naming the null before quoting the p.
Wasserman compresses the whole zoo into a chapter and is honest about which assumptions are load-bearing; Rice spends longer on the derivations and the conditions; and the likelihood-ratio chapter of any mathematical statistics text is where Wilks' theorem is actually proved. For the practical question of which test to reach for, the permutation and bootstrap material is the modern escape hatch when a shape does not fit.
- Larry Wasserman, All of Statistics: A Concise Course in Statistical Inference, chapters 10–11 — hypothesis testing and the classical tests.
- John Rice, Mathematical Statistics and Data Analysis, chapters 9–12 — the t, chi-square and F tests with their assumptions.
- George Casella and Roger Berger, Statistical Inference, chapter 10 — the likelihood-ratio test and Wilks' theorem.
- Bradley Efron and Robert Tibshirani, An Introduction to the Bootstrap, chapter 15 — permutation tests as the assumption-light alternative.
Cheat sheet: which question, which test
| Question | Test | Null distribution | Watch out for |
|---|---|---|---|
| Is this mean different from a value? | One-sample t (or z if σ known) | $t_{n-1}$ (or normal) | Skew and outliers at small n; independence |
| Do two groups differ? | Two-sample t; Welch if variances differ | $t_{n_1+n_2-2}$ (pooled) or Welch df | Paired data need a paired test, not two samples |
| Is a proportion different? | One-sample z (binomial exact for tiny n) | Normal | Small counts; $n\hat p$ and $n(1-\hat p)$ too low |
| Are two categorical variables associated? | Chi-square test of independence | $\chi^2_{(R-1)(C-1)}$ | Expected counts below five; independence of events |
| Do three or more group means differ? | One-way ANOVA (F test) | $F_{k-1,\,N-k}$ | Equal variance; a significant F does not say which group |
| Is a bigger model worth it? | Likelihood-ratio test; F for nested regressions | $\chi^2_{p}$ or $F_{p,\,N-k}$ | Extra parameters; regularity conditions for Wilks |
| Neither shape fits? | Permutation or bootstrap test | Built by resampling the data | Exchangeability under the null |