Linear regression as estimation
Fitting a line is the first thing anyone does with two variables, and it is easy to read the result as arithmetic: run the formula, get a slope, be done. But the coefficients that come out are estimates. Change the data by a hair and the line tilts; collect another sample and get another line. This part puts the whole regression report on screen at once — the fit, the standard errors of its coefficients, the residuals and a confidence band for the mean response — and then lets you drag the data to watch all four move together. The line is not a fact about the world. It is a random variable, computed from a finite sample, with a sampling distribution and an error bar.
The question
What is a fitted line, really?
You have pairs of numbers: an input $x_i$ and an output $y_i$. The scatter suggests a linear trend, so you propose a model in which the output is a straight-line function of the input plus noise,
where the errors $\varepsilon_i$ are whatever the straight line cannot explain. The two numbers $\beta_0$ and $\beta_1$ — intercept and slope — are the parameters. They describe the process that generated the data, and like every parameter in this series they are fixed and unknown. You never see them. You see ten noisy points, and you compute an estimate of them.
The estimate is the pair $(\hat\beta_0, \hat\beta_1)$ returned by ordinary least squares. It is a function of the data, so before the data arrive it is a random variable, and after they arrive it is one draw from that variable's sampling distribution. A second experiment would give a second line. This is the parameter/estimate split of the opening part reappearing in two dimensions, and the machinery of maximum likelihood and the bootstrap applies without change.
Regression is special because one solve yields a whole report: standard errors for the coefficients, a prediction interval for the mean response at any input, and a residual for every point.
The coefficients are estimates
Least squares and Gaussian maximum likelihood agree
Write the observed data as a vector $\mathbf{y}$ and build the design matrix $X$ with a leading column of ones for the intercept and a second column holding the inputs:
Least squares chooses the coefficients that make $X\boldsymbol{\beta}$ closest to $\mathbf{y}$ — the projection problem of the linear-algebra guide — with the closed-form answer
Nothing about that formula mentions probability, but add one assumption and it acquires a statistical meaning. Suppose the errors are independent Gaussians with mean zero and common variance $\sigma^2$. Then the likelihood of the data is maximised exactly where the sum of squared residuals is minimised, so the least-squares estimate is the maximum-likelihood estimate. Projection and a Gaussian likelihood select the same line, which is much of why least squares is everywhere.
That probabilistic reading turns the formula into an estimator. Because $\mathbf{y}$ is random, $\hat{\boldsymbol{\beta}}$ is a linear function of random values, hence a random vector; under the Gaussian assumption it is normal, centred on the true $\boldsymbol{\beta}$, with covariance
The diagonal entries are the coefficient variances; their square roots are the standard errors. They depend on the geometry of the inputs and on the unknown noise level $\sigma^2$, estimated by $\hat\sigma^2 = \mathrm{SSE}/(n-2)$ — the residual sum of squares over the residual degrees of freedom. Two coefficients, two degrees of freedom spent, hence $n-2$.
Stats.est.ols(X, y) takes the design matrix with its leading ones and returns the coefficients, their standard errors, the covariance matrix, the residuals, $\hat\sigma^2$, $R^2$ and the degrees of freedom, all from one normal-equation solve.
Two features of that covariance are worth noticing. The standard errors scale with $\hat\sigma$, so noisier data means wider error bars. They also scale with $(X^{\mathsf T}X)^{-1}$, which grows when the inputs are clustered or far from their mean: spread the $x$-values out and the slope is well determined, bunch them up and the line wobbles freely.
Drag the points, watch the fit
One dataset, four moving parts
The upper canvas holds the scatter, the fitted line, a 95% confidence band for the mean response, and an error bar at each input. The lower canvas holds the residuals. Grab any point: the fit recomputes, the coefficients jump, their standard errors change, the band breathes, and every residual stem updates at once.
Top: data, fitted line, 95% confidence band for the mean response, and an error bar at each input. Bottom: residuals, with the zero line for reference. Drag the round handles.
Try three experiments. Pull one point straight up and watch the line chase it while every residual changes sign; a lone observation far from the trend can move the slope because a squared error gives it enormous weight. Drag the end points outward and both standard errors fall — leverage at the extremes determines the slope best. Squeeze the cloud into a narrow vertical strip and the standard errors explode: with inputs nearly equal the design is almost singular, and the line cannot tell slope from intercept.
The residual panel is the diagnostic. Because the line forces the residuals to sum to zero and be orthogonal to the inputs, a correct plot looks centred and trend-free. What you look for is structure: a curve means the relationship bends; a widening fan means non-constant noise; one stem dwarfing the rest means a single point drives the fit. None of that appears in the coefficient readout, and all of it should stop you.
Standard errors and the confidence band
Uncertainty as a region, not a pair of numbers
The standard errors describe the wobble of the two coefficients separately. What you usually want is the uncertainty of the line: at a given input $x_0$, how far could the mean response plausibly be from the fitted value? The estimated mean is $\hat\mu(x_0) = \hat\beta_0 + \hat\beta_1 x_0$, a linear combination of the random coefficients with weights $1$ and $x_0$, so its variance is the quadratic form
Take the square root and the standard error varies with $x_0$. Two of them either side of the fitted value gives the band, and replacing $2$ with the $t$ quantile at $n-2$ degrees of freedom gives the 95% interval — the shaded region, drawn as two curves filled between. Because the variance is quadratic, the band is narrowest above the mean input $\bar x$ and flares outward with distance from the data, widening without bound beyond the observed range: the picture behind the warning against extrapolating a line.
The band's overall height is set by $\hat\sigma$, the typical residual size; its curvature is set by how the inputs are arranged — clustered inputs give a sharply pinched band that opens quickly, well-spread inputs a flatter one. Reading a regression plot is largely reading those two features.
Be precise about what it covers. It is an interval for the mean response at each $x_0$: across repeated experiments, 95% of such intervals would contain the true average output there. It is not an interval for a single new observation, which is wider because it adds that observation's own noise. The coefficient standard errors are the band in disguise: $\hat\beta_1 \pm t\,\mathrm{SE}(\hat\beta_1)$ is the range of slopes consistent with the data, and whether it contains zero is the question of whether the input does anything at all.
Residuals, $R^2$, and what they hide
The number that is not a grade
The residual is the vertical gap $e_i = y_i - \hat y_i$. Least squares forces them to sum to zero and to be orthogonal to the input column — the projection geometry of the least-squares construction in data coordinates, and why a correct residual plot always looks balanced. Their sum of squares, $\mathrm{SSE} = \sum e_i^2$, is the variation the line failed to reproduce; compared with the total variation $\mathrm{SST} = \sum (y_i - \bar y)^2$ it gives the fraction the model did reproduce:
$R^2$ is a useful one-line summary and the most over-interpreted number in applied statistics. A high $R^2$ says only that the points lie near some straight line. It does not say the relationship is linear — a U-shaped cloud can hug a line closely if its arms are short — nor that the coefficients are precise. Nearly collinear inputs can give a large $R^2$ with a slope whose confidence interval straddles zero: the fit is excellent while being unable to decide which input deserves the credit.
It is also not a mark out of a hundred to be raised by adding columns. Adding any input can only push $R^2$ up, even a column of noise, because the fit gains another way to chase the data. That is why serious reports quote an adjusted figure and, more importantly, a residual plot, which cannot be gamed by adding parameters.
The most dangerous pattern is a lone point far from the trend. Least squares squares the gap, so a point twice as far off contributes four times as much, and one at the edge of the input range has leverage to swing the line. The defence is not to delete it reflexively — it may be the most informative observation you have — but to notice it, fit with and without it, and report how much the conclusion moves.
A fitted model is a summary, not a verdict: every number you quote should be read as an estimate with a spread rather than as a fact.
Where this shows up
The same normal equation, two worlds
Regression is the first model most people fit and the last one they outgrow. Its ingredients — a design matrix, a noisy target, a squared-error criterion, a covariance matrix for the answer — reappear wherever a few numbers are estimated from many measurements, and $\hat\beta$ stays a random variable throughout.
The same fit as a projection
The normal equation $X^{\mathsf T}X\hat{\boldsymbol{\beta}} = X^{\mathsf T}\mathbf{y}$ says the residual is orthogonal to the column space of $X$, and the fitted values are the projection of $\mathbf{y}$ onto it. The least-squares part draws that geometry; this part adds the statistics on top, turning the projection into an estimator with a sampling distribution. The numerical caution about forming $X^{\mathsf T}X$ lives in the numerics part.
Calibration is weighted regression
Fitting camera intrinsics, a distortion curve or a sensor alignment to noisy measurements is least squares in a handful of parameters. The calibration part builds exactly this system and reads its residuals to judge the model — the diagnostic habit this part is built around. When the parameters are fitted by stepping rather than solving, the gradient of the same sum of squares drives the iteration, which is where the gradient and gradient descent enter.
The pattern generalises in two directions. Richer columns — polynomials, interactions, basis functions — turn the same solve into polynomial or spline regression with the covariance formula unchanged. Penalising the coefficients gives ridge regression, the ancestor of the weight decay used to train large models, trading a little bias for a large reduction in variance. Replace the closed-form solve with iteration and the linear model becomes the last layer of a network. The questions stay the same: which coefficients, how uncertain, and do the residuals look like noise.
Further reading
The references below treat regression as an estimator with a sampling distribution, which is what makes the standard errors and the confidence band meaningful rather than decorative. Wasserman covers the linear model compactly and ties it to the projection picture; Casella and Berger derive the distribution of the least-squares estimator from the Gaussian assumption.
- Larry Wasserman, All of Statistics: A Concise Course in Statistical Inference, chapter 13 — linear and logistic regression, and the geometry of least squares.
- George Casella and Roger Berger, Statistical Inference, chapter 11 — the linear model and the distribution of the least-squares estimators.
- Gareth James, Daniela Witten, Trevor Hastie and Robert Tibshirani, An Introduction to Statistical Learning, chapter 3 — linear regression with the residual plots and diagnostics this part mirrors.
- Gilbert Strang, 18.06 Linear Algebra, MIT OpenCourseWare — the projection and normal-equation lectures behind the design-matrix formula.
- 3Blue1Brown, "Essence of Linear Algebra" — the projection chapter that makes the least-squares picture geometric.
Cheat sheet
| Term | Meaning here |
|---|---|
| $\hat\beta_0, \hat\beta_1$ | Least-squares estimates of intercept and slope; random variables before the data arrive |
| $\operatorname{Cov}(\hat{\boldsymbol{\beta}})$ | $\hat\sigma^2 (X^{\mathsf T}X)^{-1}$ — the sampling covariance of the coefficients |
| Standard error | Square root of a diagonal entry of that covariance; grows with noise, shrinks with spread inputs |
| Residual $e_i$ | $y_i - \hat y_i$; sums to zero and is orthogonal to the inputs by construction |
| $\hat\sigma^2$ | $\mathrm{SSE}/(n-2)$ — the estimated noise variance, with two degrees of freedom spent on the line |
| Confidence band | $\hat\mu(x_0) \pm t_{n-2}\,\mathrm{SE}(\hat\mu(x_0))$; narrowest at $\bar x$, flares outward |
| $R^2$ | Fraction of output variance explained; high values do not imply a correct or precise model |