Math
The machinery every other subject assumes, built from pictures first: a 22-part interactive guide to linear algebra, two interactive calculus volumes that build single- and multivariable calculus from local linearity up to the gradients, Jacobians, backprop and Lie-group updates the AI and vision guides consume, two interactive probability volumes that go from what a probability means to the filters, samplers and estimators those guides run on, and an interactive guide to statistics that starts at a sample you can redraw and ends at Kalman filters and causal inference.
Linear Algebra, Interactively
An arrow, a list of numbers, and an abstract object you can add and scale — and why they are the same thing.
An arrow, a list of numbers, and an abstract object you can add and scale — and why they are the same thing.
When two vectors reach the whole plane, when they collapse to a line, and what a basis is for.
The one idea that makes matrices inevitable: a linear map is determined by where it sends the basis vectors.
Applying one transformation after another is a matrix product — which is why AB is not BA.
The signed area of the unit square after a transformation, and why det = 0 means information is destroyed.
Elimination as a sequence of row operations, and the geometric picture of three planes that meet at a point, a line, or nowhere.
Column space, null space, row space and left null space — the four subspaces that rank ties together.
Projection, cosine similarity, and the duality that a 1×n matrix is really a vector in disguise.
Turning any set of vectors into a perpendicular one, and why QR is what you actually compute.
When Ax = b has no solution, the best you can do is project b onto the column space.
The area of a parallelogram, the normal to a plane, and the matrix form of the cross product.
The same linear map looks different in every basis; P⁻¹AP is the change of costume.
Sweep a vector around a circle: almost all of them turn, except the special directions that only stretch.
Iterating x ← Ax, and why the eigenvalue magnitude decides whether a system settles, cycles or explodes.
Matrices you can rotate to face you: real eigenvalues, orthogonal eigenvectors, and xᵀAx as a shape.
Every matrix is a rotation, a stretch, and another rotation — the geometry behind everything in Act IV.
Keeping only the biggest singular values, and what PCA really does to a point cloud.
One widget for every shape of Ax = b: unique, least-squares, least-norm, and the big picture that unifies them.
Why a tiny change in b can move x a long way, and what each decomposition costs.
Rotations as matrices, angle-axis and the exponential map, gimbal lock, and why quaternions exist.
Jacobians, Hessians and the Gauss–Newton step, with the sparsity pattern that makes pose graphs tractable.
Attention as three matrices, and a low-rank update to a weight matrix — the linear algebra a transformer runs on.
Every term in the series linked back to where it is introduced, plus a cheat sheet of decompositions and identities.
Calculus, Interactively
The whole subject in one idea: almost every function is locally linear, and a derivative is just the slope you find when you zoom in far enough.
The whole subject in one idea: almost every function is locally linear, and a derivative is just the slope you find when you zoom in far enough.
The epsilon-delta definition as a two-slider game, one-sided limits, and the exact places a function can fail to be continuous.
A secant line sliding into a tangent: drag the point and watch the derivative trace itself out, then see why the slope is the best linear approximation.
Product, quotient and chain rules built as pictures - the product rule as the area of a growing rectangle - plus the derivatives worth memorising.
Composed machines with rate gauges: multiply the gear ratios and the chain rule falls out, the same rule backprop runs in reverse.
Differentiate an equation you cannot solve for y by dragging along the curve it defines, and connect rates that are tied together by a constraint.
Add terms one at a time and watch a polynomial wrap a curve, then use the first-order model to run Newton's method and map its basins of attraction.
A Riemann-sum slider that shrinks rectangles into signed area, with left, right, midpoint and trapezoid rules side by side.
Area-so-far traced live beside the curve: differentiation and integration are inverse operations, and the odometer makes the theorem an observation.
Substitution as the chain rule backwards, integration by parts as the product rule backwards, and when to give up on closed forms and use quadrature.
A surface with draggable slice planes: hold one input fixed and single-variable calculus returns as a partial derivative.
A contour map, a directional-derivative dial, and the reason the gradient points uphill and meets every level set at a right angle.
The arm Jacobian, a quadratic-form explorer, and Hessian eigenvalues that sort every critical point into bowl, ridge or saddle.
The local quadratic model on a surface, the second-derivative test in matrix form, and the shape of the landscape a solver has to descend.
Every notation and identity used across the volume, collected in one place with a pointer back to the part that derives it.
Calculus in Motion, Interactively
A particle-flow canvas you edit by hand: arrows at every point, streamlines through them, and the two numbers that describe how a field changes.
A particle-flow canvas you edit by hand: arrows at every point, streamlines through them, and the two numbers that describe how a field changes.
Drag a path through a field and compare the work done; when the answer depends only on the endpoints, a potential surface is hiding underneath.
Drop a shrinking box probe and a paddlewheel into a field you shape, and read off outflow and rotation point by point.
Grow and shrink a region and watch the boundary total and the interior total stay equal - the Fundamental Theorem wearing three different costumes.
Deform a grid and watch cell areas scale by the determinant, the factor that makes a change of variables exact.
Double and triple integrals swept in polar, cylindrical and spherical coordinates, with the volume element each one brings.
Drag a constraint against the level sets of an objective and watch tangency appear as the condition for a constrained optimum, KKT inequalities included.
Layout conventions, a shape checker, and the handful of matrix-derivative identities that do most of the work in machine learning.
Build a small computation graph, run it forward, and watch the adjoints fill in backwards - the chain rule as an algorithm.
Race gradient descent, momentum and Adam on a surface you shape, and see how the condition number explains every zig-zag and slow crawl.
Push a Gaussian through a map, follow how densities transform with the Jacobian, and meet reparameterisation and the KL divergence.
A slope field you drop solutions into, Euler versus RK4 versus exact, and the pendulum and unicycle that turn into stability questions.
A naive update versus an exponential-map update on a sphere, and why rotations do not add - differentiation where the space is curved.
Drag two endpoints and watch the minimum-jerk path re-solve itself: optimising over whole functions, and the route to trajectory optimisation.
A routing table from every idea in the volume back into the AI, vision and robotics guides that consume it.
Probability, Interactively
Relative frequency settling into a limit, a belief you can bet on, and the three axioms both pictures have to obey.
Relative frequency settling into a limit, a belief you can bet on, and the three axioms both pictures have to obey.
Permutations against combinations, the birthday problem, and the hash collisions that make counting a practical tool.
Conditioning literally crops the sample space and renormalizes it — the unit square squeezed into a corner.
The medical test, natural frequencies, and the odds form that lets evidence accumulate one log at a time.
Conditional independence, explaining-away in a three-node network, and Simpson's paradox in two draggable clouds.
One sample, three views: outcomes to PMF bars to a CDF staircase, with the highlighting linked across all three.
The balance point of a distribution, why linearity survives dependence, LOTUS, and the indicator trick.
Variance as a moment of inertia, and Markov, Chebyshev and Hoeffding bounds drawn over the true tail.
Bernoulli, binomial, geometric, negative binomial and Poisson in one switchable panel, with the Poisson limit animated.
Narrow the bins until a histogram becomes a pdf, and meet the trap that a density is not a probability.
Uniform, exponential, gamma, beta and Student-t, and the Gaussian as the distribution entropy and the CLT both select.
Push a density through a monotone map, watch it squash by the Jacobian, and sample by inverting the CDF.
A joint table with its marginals on the edges: slicing a row is conditioning, and independence is an outer product.
Drag a point cloud and read off the covariance matrix, the correlation, and the 1σ and 2σ ellipses it defines.
Eigen-decomposition as the ellipse axes, Mahalanobis distance, live conditioning, whitening, and a 3D surface.
Slide one density across another to convolve them, and see why the Gaussian is stable while the Cauchy is not.
Many running averages at once, weak against strong, and the four modes of convergence as a map of counterexamples.
Pick a base distribution, slide n, and watch the sample mean become Gaussian — with the Berry–Esseen error readout.
Where the CLT fails: a Cauchy mean that never settles, Chernoff bounds, and concentration of measure in 100 dimensions.
Why not every set can have a probability, σ-algebras as the questions you may ask, filtrations, and Borel–Cantelli.
Every term in the volume linked back to the part that introduces it, plus a reference card of pmfs, pdfs, means, variances and mgfs.
Probability in Action, Interactively
Drag the parameter and watch the log-likelihood surface move with the fitted curve — and see why a likelihood is not a probability over θ.
Drag the parameter and watch the log-likelihood surface move with the fitted curve — and see why a likelihood is not a probability over θ.
Shape a Beta prior by hand, stream in data a point at a time, and watch the posterior concentrate.
Sampling distributions, the bias–variance trade-off on a dartboard, Fisher information, and the Cramér–Rao floor.
A hundred simulated intervals with about ninety-five covering the truth, beside the posterior interval on the same data.
Null against alternative with a draggable effect size: α, β and power, the p-value under H₀, and p-hacking made visible.
Resample with replacement and build a sampling distribution from one dataset — and see where the method breaks.
Inverse-CDF, rejection sampling with a draggable envelope, and importance sampling with its weight histogram.
π by darts with its 1/√N error band, then antithetic variates, control variates and stratification raced on one plot.
Walks in one and two dimensions, √n scaling, exponential gaps, the Brownian limit, and a martingale with optional stopping.
Edit a transition matrix, watch a token hop between states, and see the stationary distribution appear — then rank the web.
Metropolis–Hastings on a 2D posterior, the step size that helps and then hurts, Gibbs' axis-aligned moves, and HMC for contrast.
Fit a Gaussian mixture step by step, with responsibilities as colour and the ELBO as a lower bound you can watch tighten.
Shape a distribution and watch entropy fall, read KL as a codelength excess, and see mode-seeking against mode-covering fits.
A clickable DAG with d-separation highlighted, explaining-away done numerically, and the same model drawn as a factor graph.
Forward–backward fills a trellis cell by cell and Viterbi's backtrace lights up, on a robot lost in a corridor.
A histogram filter on a looped corridor: predict is convolve and blur, update is multiply and sharpen — the loop everything else specialises.
The same recursion with Gaussians: a tunable tracker, then a 2D EKF whose linearisation error you can watch grow.
Monte Carlo localization in a 2D map — sample, weight, resample — with particle depletion shown and then fixed.
Heavy-tailed noise breaking least squares, RANSAC's inlier search, and the Huber and Cauchy loss curves that repair it.
Softmax as a distribution with a temperature dial, cross-entropy as negative log-likelihood, calibration, a GP posterior, VAEs and diffusion.
A routing table from every idea in the volume back into the AI, vision and robotics guides that consume it.
Statistics, Interactively
A hidden population, one sample, and an estimate that moves every time you redraw it while the truth stays put.
A hidden population, one sample, and an estimate that moves every time you redraw it while the truth stays put.
Histogram, KDE and ECDF of the same data, and how a single outlier pulls the mean away from the median.
Two variables, one ellipse: drag the correlation and the points, and meet the four datasets with identical summaries.
Repeatedly draw a sample, drop its mean into a growing histogram, and watch the spread shrink as the sample grows.
Convolve a draggable population with itself, and read the standard error off a curve instead of a formula.
Race three estimators of the same quantity across thousands of simulated samples, and split the error into the part that shrinks and the part that does not.
Data as ticks on an axis; drag the parameters over a live log-likelihood surface and watch the density slide onto the data.
The curvature of the log-likelihood is the error bar: zoom the peak, read off the width, and compare it to the best any estimator can do.
Resample the data in front of you, build the bootstrap distribution, and get an interval without a formula.
Drag points and watch the fit, its standard error, the residuals and the confidence band move together.
A hundred repeated experiments, a hundred intervals, and a count of how many miss. The interval is random; the parameter is not.
Shuffle the labels, rebuild the null distribution, and shade the tail - a permutation test you can watch, with no formula required.
Two overlapping densities: drag the effect size, the sample size and the threshold, and watch the four outcomes trade against each other.
A p-hacking sandbox: add covariates and subgroups until something is significant, then correct with Bonferroni or Benjamini-Hochberg.
One interactive reference: pick a question, and see which test, which null distribution and which assumption it needs.
Beta-Binomial inference: drag the prior, feed coin flips one at a time, and watch the data overwhelm it.
Two Gaussians merging into a precision-weighted average - deliberately the seed of the Kalman filter.
Credible vs confidence side by side, and a loss function that decides whether the mean, the median or the mode is the number to report.
Live Metropolis-Hastings on a two-dimensional posterior: proposal width, trace plot, acceptance rate, burn-in and R-hat.
Fit a Gaussian to a skewed posterior, and toggle KL(q‖p) against KL(p‖q) to see mode-seeking and mass-covering.
A polynomial-degree slider and an ensemble of fits over resampled data, drawing the bias and variance terms directly.
Ridge is a Gaussian prior and lasso is a Laplace prior: watch the constraint region and the loss contours meet, and see MAP become penalised maximum likelihood.
A k-fold animation, and the optimism of training error made visible as a gap that grows with capacity.
One threshold slider driving the confusion matrix, the ROC point, the PR point and the reliability diagram together.
Ensembles, MC dropout and split conformal prediction: the coverage guarantee demonstrated, then broken by distribution shift.
A draggable covariance, marginals on the margins and a conditioning slice: the multivariate Gaussian made geometric.
A robot in a corridor: a histogram belief that blurs when it moves and multiplies when it sees a door.
Track in 2D with a live covariance ellipse, tune Q and R until it diverges, then linearise a range-bearing model and watch the EKF's error.
Particles over a map, systematic resampling, effective sample size, particle deprivation and the kidnapped-robot problem.
A sequential-testing simulator: peek early and often, and watch the false-positive rate climb far past the level you chose.
An interactive DAG: fork, chain and collider. Condition on a collider and watch a spurious correlation appear from nothing.
Potential outcomes on one simulated dataset: matching, inverse-probability weighting and difference-in-differences, and the estimate moving as you adjust.
Every term in the series linked back to where it is introduced, plus the distribution table, a test-selection table and the estimator and interval formula card.
Why this section exists
Every other subject here leans on linear algebra, calculus, probability and statistics it never teaches. The multi-view geometry guide factorises an essential matrix and reasons about RANSAC inlier probabilities, the nonlinear optimization guide solves normal equations and takes Newton steps, and the LLM guides are matrix multiplies, gradients, cross-entropy and calibration all the way down. This section is where that machinery gets built properly — geometry first, algebra second, every idea attached to something you can drag.
Linear Algebra, Interactively starts at what a vector is and ends at attention matrices, Gauss-Newton and SO(3). Calculus, Interactively is single-variable calculus first, carried by a robot moving along a line, then multivariable calculus, carried by a two-link arm whose Jacobian is the derivative as a matrix; Calculus in Motion lets that arm move — vector fields, constrained optimization, matrix calculus and backprop, differential equations, and calculus on manifolds. Neither calculus volume re-teaches gradient descent, Newton's method or the matrix exponential as algorithms — those live in Nonlinear Optimization, which these parts feed by name. Every part in every series ends with a "where this shows up" line pointing at the page elsewhere on the site that uses it.
Probability, Interactively begins at what a probability means and works up through counting, conditioning, Bayes' rule, random variables and the limit theorems. Probability in Action takes those foundations and runs them: likelihood and Bayesian inference, estimators and intervals, sampling and Monte Carlo, Markov chains and MCMC, information theory, and finally the Bayes, Kalman and particle filters that SLAM is built from, plus the softmax, cross-entropy and calibration the language-model guides assume. The change-of-variables calculus those models need is in Calculus in Motion; here we do the probability of it. Statistics, Interactively is the newest volume: 32 parts from a single sample you can redraw through estimation, testing and Bayesian inference to the Kalman and particle filters the robotics guides run, ending at causal inference. It assumes probability as a prerequisite — sample spaces, conditional probability, random variables, expectation and the named distributions, plus the statements of the law of large numbers and the central limit theorem — and Part 1 recaps only the load-bearing pieces.