Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

Linear Algebra, Interactively

An arrow, a list of numbers, and an abstract object you can add and scale — and why they are the same thing.

Part 1
What a vector is (three answers)

An arrow, a list of numbers, and an abstract object you can add and scale — and why they are the same thing.

Part 2
Span, independence and basis

When two vectors reach the whole plane, when they collapse to a line, and what a basis is for.

Part 3
Linear transformations are matrices

The one idea that makes matrices inevitable: a linear map is determined by where it sends the basis vectors.

Part 4
Matrix multiplication is composition

Applying one transformation after another is a matrix product — which is why AB is not BA.

Part 5
The determinant: signed area

The signed area of the unit square after a transformation, and why det = 0 means information is destroyed.

Part 6
Solving Ax = b, and Gaussian elimination

Elimination as a sequence of row operations, and the geometric picture of three planes that meet at a point, a line, or nowhere.

Part 7
Rank, null space, and the four subspaces

Column space, null space, row space and left null space — the four subspaces that rank ties together.

Part 8
Dot product, projection, duality

Projection, cosine similarity, and the duality that a 1×n matrix is really a vector in disguise.

Part 9
Orthonormal bases and Gram–Schmidt

Turning any set of vectors into a perpendicular one, and why QR is what you actually compute.

Part 10
Least squares is a projection

When Ax = b has no solution, the best you can do is project b onto the column space.

Part 11
Cross product and skew-symmetric matrices

The area of a parallelogram, the normal to a plane, and the matrix form of the cross product.

Part 12
Same map, different coordinates

The same linear map looks different in every basis; P⁻¹AP is the change of costume.

Part 13
The arrows that don't turn

Sweep a vector around a circle: almost all of them turn, except the special directions that only stretch.

Part 14
Matrix powers and dynamical systems

Iterating x ← Ax, and why the eigenvalue magnitude decides whether a system settles, cycles or explodes.

Part 15
Symmetric matrices and quadratic forms

Matrices you can rotate to face you: real eigenvalues, orthogonal eigenvectors, and xᵀAx as a shape.

Part 16
Every matrix maps a circle to an ellipse

Every matrix is a rotation, a stretch, and another rotation — the geometry behind everything in Act IV.

Part 17
Low-rank approximation and PCA

Keeping only the biggest singular values, and what PCA really does to a point cloud.

Part 18
The complete picture of Ax = b

One widget for every shape of Ax = b: unique, least-squares, least-norm, and the big picture that unifies them.

Part 19
Conditioning, stability, and cost

Why a tiny change in b can move x a long way, and what each decomposition costs.

Part 20
Rotations, SO(3) and the exponential map

Rotations as matrices, angle-axis and the exponential map, gimbal lock, and why quaternions exist.

Part 21
Jacobians, Hessians and Gauss–Newton

Jacobians, Hessians and the Gauss–Newton step, with the sparsity pattern that makes pose graphs tractable.

Part 22
Linear algebra in ML and AI

Attention as three matrices, and a low-rank update to a weight matrix — the linear algebra a transformer runs on.

Reference
Glossary and identities to know

Every term in the series linked back to where it is introduced, plus a cheat sheet of decompositions and identities.

Open the 22-part guide →

Calculus, Interactively

The whole subject in one idea: almost every function is locally linear, and a derivative is just the slope you find when you zoom in far enough.

Part 1
Zoom in until it's a line

The whole subject in one idea: almost every function is locally linear, and a derivative is just the slope you find when you zoom in far enough.

Part 2
Limits, made concrete

The epsilon-delta definition as a two-slider game, one-sided limits, and the exact places a function can fail to be continuous.

Part 3
The derivative

A secant line sliding into a tangent: drag the point and watch the derivative trace itself out, then see why the slope is the best linear approximation.

Part 4
A toolbox of derivatives

Product, quotient and chain rules built as pictures - the product rule as the area of a growing rectangle - plus the derivatives worth memorising.

Part 5
The chain rule

Composed machines with rate gauges: multiply the gear ratios and the chain rule falls out, the same rule backprop runs in reverse.

Part 6
Implicit differentiation & related rates

Differentiate an equation you cannot solve for y by dragging along the curve it defines, and connect rates that are tied together by a constraint.

Part 7
Linear approximation, Newton, Taylor

Add terms one at a time and watch a polynomial wrap a curve, then use the first-order model to run Newton's method and map its basins of attraction.

Part 8
The integral as accumulation

A Riemann-sum slider that shrinks rectangles into signed area, with left, right, midpoint and trapezoid rules side by side.

Part 9
The Fundamental Theorem

Area-so-far traced live beside the curve: differentiation and integration are inverse operations, and the odometer makes the theorem an observation.

Part 10
Substitution, parts, and going numeric

Substitution as the chain rule backwards, integration by parts as the product rule backwards, and when to give up on closed forms and use quadrature.

Part 11
Functions of several variables

A surface with draggable slice planes: hold one input fixed and single-variable calculus returns as a partial derivative.

Part 12
The gradient

A contour map, a directional-derivative dial, and the reason the gradient points uphill and meets every level set at a right angle.

Part 13
The derivative as a matrix

The arm Jacobian, a quadratic-form explorer, and Hessian eigenvalues that sort every critical point into bowl, ridge or saddle.

Part 14
Multivariable Taylor & critical points

The local quadratic model on a surface, the second-derivative test in matrix form, and the shape of the landscape a solver has to descend.

Glossary
Notation & identities to know

Every notation and identity used across the volume, collected in one place with a pointer back to the part that derives it.

Open the 14-part guide →

Calculus in Motion, Interactively

A particle-flow canvas you edit by hand: arrows at every point, streamlines through them, and the two numbers that describe how a field changes.

Part 1
Vector fields

A particle-flow canvas you edit by hand: arrows at every point, streamlines through them, and the two numbers that describe how a field changes.

Part 2
Line integrals & conservative fields

Drag a path through a field and compare the work done; when the answer depends only on the endpoints, a potential surface is hiding underneath.

Part 3
Divergence and curl

Drop a shrinking box probe and a paddlewheel into a field you shape, and read off outflow and rotation point by point.

Part 4
Green, Stokes, Divergence: one theorem

Grow and shrink a region and watch the boundary total and the interior total stay equal - the Fundamental Theorem wearing three different costumes.

Part 5
The Jacobian determinant

Deform a grid and watch cell areas scale by the determinant, the factor that makes a change of variables exact.

Part 6
Multiple integrals & coordinates

Double and triple integrals swept in polar, cylindrical and spherical coordinates, with the volume element each one brings.

Part 7
Lagrange multipliers & KKT

Drag a constraint against the level sets of an objective and watch tangency appear as the condition for a constrained optimum, KKT inequalities included.

Part 8
Matrix calculus

Layout conventions, a shape checker, and the handful of matrix-derivative identities that do most of the work in machine learning.

Part 9
Backpropagation is the chain rule

Build a small computation graph, run it forward, and watch the adjoints fill in backwards - the chain rule as an algorithm.

Part 10
Curvature and convergence

Race gradient descent, momentum and Adam on a surface you shape, and see how the condition number explains every zig-zag and slow crawl.

Part 11
Calculus of probability

Push a Gaussian through a map, follow how densities transform with the Jacobian, and meet reparameterisation and the KL divergence.

Part 12
Differential equations

A slope field you drop solutions into, Euler versus RK4 versus exact, and the pendulum and unicycle that turn into stability questions.

Part 13
Calculus on manifolds

A naive update versus an exponential-map update on a sphere, and why rotations do not add - differentiation where the space is curved.

Part 14
Calculus of variations & optimal control

Drag two endpoints and watch the minimum-jerk path re-solve itself: optimising over whole functions, and the route to trajectory optimisation.

Glossary
Where each idea is used on this site

A routing table from every idea in the volume back into the AI, vision and robotics guides that consume it.

Open the 14-part guide →

Probability, Interactively

Relative frequency settling into a limit, a belief you can bet on, and the three axioms both pictures have to obey.

Part 1
What a probability is

Relative frequency settling into a limit, a belief you can bet on, and the three axioms both pictures have to obey.

Part 2
Counting, and why it is the hard part

Permutations against combinations, the birthday problem, and the hash collisions that make counting a practical tool.

Part 3
Conditional probability

Conditioning literally crops the sample space and renormalizes it — the unit square squeezed into a corner.

Part 4
Bayes' rule

The medical test, natural frequencies, and the odds form that lets evidence accumulate one log at a time.

Part 5
Independence, and the paradoxes

Conditional independence, explaining-away in a three-node network, and Simpson's paradox in two draggable clouds.

Part 6
Random variables, PMFs and CDFs

One sample, three views: outcomes to PMF bars to a CDF staircase, with the highlighting linked across all three.

Part 7
Expectation

The balance point of a distribution, why linearity survives dependence, LOTUS, and the indicator trick.

Part 8
Spread, moments and concentration

Variance as a moment of inertia, and Markov, Chebyshev and Hoeffding bounds drawn over the true tail.

Part 9
The discrete families

Bernoulli, binomial, geometric, negative binomial and Poisson in one switchable panel, with the Poisson limit animated.

Part 10
From mass to density

Narrow the bins until a histogram becomes a pdf, and meet the trap that a density is not a probability.

Part 11
The continuous families and the Gaussian

Uniform, exponential, gamma, beta and Student-t, and the Gaussian as the distribution entropy and the CLT both select.

Part 12
Transforming a random variable

Push a density through a monotone map, watch it squash by the Jacobian, and sample by inverting the CDF.

Part 13
Joint, marginal, conditional

A joint table with its marginals on the edges: slicing a row is conditioning, and independence is an outer product.

Part 14
Covariance, correlation and Σ

Drag a point cloud and read off the covariance matrix, the correlation, and the 1σ and 2σ ellipses it defines.

Part 15
The multivariate Gaussian

Eigen-decomposition as the ellipse axes, Mahalanobis distance, live conditioning, whitening, and a 3D surface.

Part 16
Sums, convolution and generating functions

Slide one density across another to convolve them, and see why the Gaussian is stable while the Cauchy is not.

Part 17
The Law of Large Numbers

Many running averages at once, weak against strong, and the four modes of convergence as a map of counterexamples.

Part 18
The Central Limit Theorem

Pick a base distribution, slide n, and watch the sample mean become Gaussian — with the Berry–Esseen error readout.

Part 19
Tails, heavy tails, and high dimensions

Where the CLT fails: a Cauchy mean that never settles, Chernoff bounds, and concentration of measure in 100 dimensions.

Part 20
A measure-theoretic aside

Why not every set can have a probability, σ-algebras as the questions you may ask, filtrations, and Borel–Cantelli.

Glossary
Glossary and distribution reference

Every term in the volume linked back to the part that introduces it, plus a reference card of pmfs, pdfs, means, variances and mgfs.

Open the 20-part guide →

Probability in Action, Interactively

Drag the parameter and watch the log-likelihood surface move with the fitted curve — and see why a likelihood is not a probability over θ.

Part 1
Likelihood and maximum likelihood

Drag the parameter and watch the log-likelihood surface move with the fitted curve — and see why a likelihood is not a probability over θ.

Part 2
Prior, posterior, conjugacy

Shape a Beta prior by hand, stream in data a point at a time, and watch the posterior concentrate.

Part 3
Is this estimator any good?

Sampling distributions, the bias–variance trade-off on a dartboard, Fisher information, and the Cramér–Rao floor.

Part 4
Confidence vs credible

A hundred simulated intervals with about ninety-five covering the truth, beside the posterior interval on the same data.

Part 5
Hypothesis testing, and how it goes wrong

Null against alternative with a draggable effect size: α, β and power, the p-value under H₀, and p-hacking made visible.

Part 6
The bootstrap

Resample with replacement and build a sampling distribution from one dataset — and see where the method breaks.

Part 7
How to draw a sample

Inverse-CDF, rejection sampling with a draggable envelope, and importance sampling with its weight histogram.

Part 8
Monte Carlo integration

π by darts with its 1/√N error band, then antithetic variates, control variates and stratification raced on one plot.

Part 9
Random walks, Poisson processes, Brownian motion

Walks in one and two dimensions, √n scaling, exponential gaps, the Brownian limit, and a martingale with optional stopping.

Part 10
Markov chains

Edit a transition matrix, watch a token hop between states, and see the stationary distribution appear — then rank the web.

Part 11
MCMC

Metropolis–Hastings on a 2D posterior, the step size that helps and then hurts, Gibbs' axis-aligned moves, and HMC for contrast.

Part 12
Latent variables: EM and the ELBO

Fit a Gaussian mixture step by step, with responsibilities as colour and the ELBO as a lower bound you can watch tighten.

Part 13
Entropy, cross-entropy, KL, mutual information

Shape a distribution and watch entropy fall, read KL as a codelength excess, and see mode-seeking against mode-covering fits.

Part 14
Bayes nets and factor graphs

A clickable DAG with d-separation highlighted, explaining-away done numerically, and the same model drawn as a factor graph.

Part 15
Hidden Markov models

Forward–backward fills a trellis cell by cell and Viterbi's backtrace lights up, on a robot lost in a corridor.

Part 16
The Bayes filter

A histogram filter on a looped corridor: predict is convolve and blur, update is multiply and sharpen — the loop everything else specialises.

Part 17
The Kalman filter and EKF

The same recursion with Gaussians: a tunable tracker, then a 2D EKF whose linearisation error you can watch grow.

Part 18
Particle filters

Monte Carlo localization in a 2D map — sample, weight, resample — with particle depletion shown and then fixed.

Part 19
Outliers, RANSAC and robust estimation

Heavy-tailed noise breaking least squares, RANSAC's inlier search, and the Huber and Cauchy loss curves that repair it.

Part 20
Probability in machine learning

Softmax as a distribution with a temperature dial, cross-entropy as negative log-likelihood, calibration, a GP posterior, VAEs and diffusion.

Glossary
Where each idea is used on this site

A routing table from every idea in the volume back into the AI, vision and robotics guides that consume it.

Open the 20-part guide →

Statistics, Interactively

A hidden population, one sample, and an estimate that moves every time you redraw it while the truth stays put.

Part 1
What statistics actually asks

A hidden population, one sample, and an estimate that moves every time you redraw it while the truth stays put.

Part 2
Describing one sample

Histogram, KDE and ECDF of the same data, and how a single outlier pulls the mean away from the median.

Part 3
Covariance, correlation, and the ellipse

Two variables, one ellipse: drag the correlation and the points, and meet the four datasets with identical summaries.

Part 4
The sampling distribution

Repeatedly draw a sample, drop its mean into a growing histogram, and watch the spread shrink as the sample grows.

Part 5
The CLT you can feel

Convolve a draggable population with itself, and read the standard error off a curve instead of a formula.

Part 6
Bias, variance, and mean squared error

Race three estimators of the same quantity across thousands of simulated samples, and split the error into the part that shrinks and the part that does not.

Part 7
Maximum likelihood

Data as ticks on an axis; drag the parameters over a live log-likelihood surface and watch the density slide onto the data.

Part 8
Fisher information and the Cramér–Rao bound

The curvature of the log-likelihood is the error bar: zoom the peak, read off the width, and compare it to the best any estimator can do.

Part 9
The bootstrap and the plug-in principle

Resample the data in front of you, build the bootstrap distribution, and get an interval without a formula.

Part 10
Linear regression as estimation

Drag points and watch the fit, its standard error, the residuals and the confidence band move together.

Part 11
What the 95% refers to

A hundred repeated experiments, a hundred intervals, and a count of how many miss. The interval is random; the parameter is not.

Part 12
Null distributions and the p-value

Shuffle the labels, rebuild the null distribution, and shade the tail - a permutation test you can watch, with no formula required.

Part 13
Errors, power, and sample size

Two overlapping densities: drag the effect size, the sample size and the threshold, and watch the four outcomes trade against each other.

Part 14
p-hacking and the garden of forking paths

A p-hacking sandbox: add covariates and subgroups until something is significant, then correct with Bonferroni or Benjamini-Hochberg.

Part 15
t, z, chi-square, F and the likelihood-ratio

One interactive reference: pick a question, and see which test, which null distribution and which assumption it needs.

Part 16
Prior × likelihood → posterior

Beta-Binomial inference: drag the prior, feed coin flips one at a time, and watch the data overwhelm it.

Part 17
Conjugacy and Gaussian fusion

Two Gaussians merging into a precision-weighted average - deliberately the seed of the Kalman filter.

Part 18
Credible intervals and decision theory

Credible vs confidence side by side, and a loss function that decides whether the mean, the median or the mode is the number to report.

Part 19
When conjugacy fails: MCMC

Live Metropolis-Hastings on a two-dimensional posterior: proposal width, trace plot, acceptance rate, burn-in and R-hat.

Part 20
Laplace and variational approximations

Fit a Gaussian to a skewed posterior, and toggle KL(q‖p) against KL(p‖q) to see mode-seeking and mass-covering.

Part 21
The bias–variance decomposition

A polynomial-degree slider and an ensemble of fits over resampled data, drawing the bias and variance terms directly.

Part 22
Regularization is a prior

Ridge is a Gaussian prior and lasso is a Laplace prior: watch the constraint region and the loss contours meet, and see MAP become penalised maximum likelihood.

Part 23
Cross-validation, AIC and BIC

A k-fold animation, and the optimism of training error made visible as a gap that grows with capacity.

Part 24
Evaluation, ROC/PR, and calibration

One threshold slider driving the confusion matrix, the ROC point, the PR point and the reliability diagram together.

Part 25
Predictive uncertainty

Ensembles, MC dropout and split conformal prediction: the coverage guarantee demonstrated, then broken by distribution shift.

Part 26
Gaussians in n dimensions

A draggable covariance, marginals on the margins and a conditioning slice: the multivariate Gaussian made geometric.

Part 27
The Bayes filter

A robot in a corridor: a histogram belief that blurs when it moves and multiplies when it sees a door.

Part 28
The Kalman filter and the EKF

Track in 2D with a live covariance ellipse, tune Q and R until it diverges, then linearise a range-bearing model and watch the EKF's error.

Part 29
Particle filters and MCL

Particles over a map, systematic resampling, effective sample size, particle deprivation and the kidnapped-robot problem.

Part 30
Randomization, A/B tests, and peeking

A sequential-testing simulator: peek early and often, and watch the false-positive rate climb far past the level you chose.

Part 31
Confounding, Simpson's paradox, and DAGs

An interactive DAG: fork, chain and collider. Condition on a collider and watch a spurious correlation appear from nothing.

Part 32
Estimating causal effects

Potential outcomes on one simulated dataset: matching, inverse-probability weighting and difference-in-differences, and the estimate moving as you adjust.

Reference
Glossary, distribution table and cheat cards

Every term in the series linked back to where it is introduced, plus the distribution table, a test-selection table and the estimator and interval formula card.

Open the 32-part guide →

Why this section exists

Every other subject here leans on linear algebra, calculus, probability and statistics it never teaches. The multi-view geometry guide factorises an essential matrix and reasons about RANSAC inlier probabilities, the nonlinear optimization guide solves normal equations and takes Newton steps, and the LLM guides are matrix multiplies, gradients, cross-entropy and calibration all the way down. This section is where that machinery gets built properly — geometry first, algebra second, every idea attached to something you can drag.

Linear Algebra, Interactively starts at what a vector is and ends at attention matrices, Gauss-Newton and SO(3). Calculus, Interactively is single-variable calculus first, carried by a robot moving along a line, then multivariable calculus, carried by a two-link arm whose Jacobian is the derivative as a matrix; Calculus in Motion lets that arm move — vector fields, constrained optimization, matrix calculus and backprop, differential equations, and calculus on manifolds. Neither calculus volume re-teaches gradient descent, Newton's method or the matrix exponential as algorithms — those live in Nonlinear Optimization, which these parts feed by name. Every part in every series ends with a "where this shows up" line pointing at the page elsewhere on the site that uses it.

Probability, Interactively begins at what a probability means and works up through counting, conditioning, Bayes' rule, random variables and the limit theorems. Probability in Action takes those foundations and runs them: likelihood and Bayesian inference, estimators and intervals, sampling and Monte Carlo, Markov chains and MCMC, information theory, and finally the Bayes, Kalman and particle filters that SLAM is built from, plus the softmax, cross-entropy and calibration the language-model guides assume. The change-of-variables calculus those models need is in Calculus in Motion; here we do the probability of it. Statistics, Interactively is the newest volume: 32 parts from a single sample you can redraw through estimation, testing and Bayesian inference to the Kalman and particle filters the robotics guides run, ending at causal inference. It assumes probability as a prerequisite — sample spaces, conditional probability, random variables, expectation and the named distributions, plus the statements of the law of large numbers and the central limit theorem — and Part 1 recaps only the load-bearing pieces.