Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

The question

Same randomness, new coordinates

A random variable is a number produced by a chance experiment. If X is that number, then Y=g(X) is also a number produced by the same experiment — you just read it differently. If X is a measured voltage in millivolts and g converts to volts, Y is the same measurement in new units. If X is the time until a component fails and g is the logarithm, Y is the failure time on a log scale. Nothing new happened physically. But the distribution changed, and predicting it is the job.

The clean way to think about it is in terms of mass. Pick a tiny interval [x,x+dx]. The probability that X lands there is $f_X(x)\,dx$. The map sends that interval to a tiny interval around y=g(x), whose width is $dy=g'(x)\,dx$. All of the probability that was in the first interval must now be in the second one, because the map moved it there and moved nothing else. Two intervals, one probability: $f_X(x)\,dx = f_Y(y)\,dy$. Divide and the density ratio is dx/dy, the reciprocal of the local stretch. Stretch the axis and the density thins; compress it and the density piles up. That is the whole idea, and everything on this page is a consequence of it.

There are two cases to keep separate. A monotone map is one that never turns around: strictly increasing or strictly decreasing, so each output value comes from exactly one input value and an inverse exists. Then the bookkeeping is a single derivative and the formula is elementary. A non-monotone map can send several inputs to the same output, and each of those inputs deposits its probability at the same place; the density picks up a sum of terms. The derivative alone is not enough. We will handle the monotone case properly first, mention the sum for the general one, and then specialise to the most useful map of all: the inverse of the cumulative distribution function.

The inverse-CDF trick is worth flagging now because it inverts the usual direction. Instead of asking what distribution comes out when a given variable goes through a given map, it asks which map turns a uniform variable into a given distribution. The answer is beautifully simple, and it is the reason every simulation library in the world can produce a Gaussian, an exponential, or a beta from a stream of uniform bits. The change-of-variables formula is what proves it correct. Jensen's inequality, at the end, is the same integral identity viewed through convexity, and it survives even when no density exists at all.

💡 By the end of this part you'll see why a density rescales by the reciprocal of the map's derivative, why inverting a CDF turns uniforms into samples from any distribution, and why a convex function can only push an expectation upward.
2

The change-of-variables formula

A density is a stretchy ruler

Let X have density f_X and let g be strictly monotone with inverse g^{-1}. The density of Y=g(X) is the input density, evaluated at the pre-image of y, multiplied by how much the inverse map stretches a unit of y:

$$f_Y(y)=f_X\!\left(g^{-1}(y)\right)\left|\frac{d}{dy}\,g^{-1}(y)\right|.$$

The absolute value is not decoration. If g is decreasing, the derivative of the inverse is negative, but a density is never negative. The map has reflected the axis as well as stretched it, and the reflection changes which end is which without changing any probability; so we take the magnitude and the reflection is handled. For an increasing g the inverse derivative is 1/g'(x) with x=g^{-1}(y), which gives the friendlier form

$$g \text{ increasing:}\quad f_Y(y)=\frac{f_X\!\left(g^{-1}(y)\right)}{g'\!\left(g^{-1}(y)\right)},\qquad g \text{ decreasing:}\quad f_Y(y)=\frac{f_X\!\left(g^{-1}(y)\right)}{\left|g'\!\left(g^{-1}(y)\right)\right|}.$$

Read the fraction as a statement about units. The numerator is a probability per unit of x; the denominator converts units of x into units of y. If g is steep, a unit of y spans only a sliver of x, so the same probability is packed into a narrower output interval and the density rises. Multiplying a variable by ten divides its density by ten. Exponentiating it produces a density that is thin at the right and heavy at the left, which is exactly the lognormal shape.

In several variables the local stretch stops being one number and becomes the determinant of the Jacobian. If $\mathbf{Y}=g(\mathbf{X})$ is a one-to-one smooth map of $\mathbb{R}^n$, then

$$f_{\mathbf{Y}}(\mathbf{y})=f_{\mathbf{X}}\!\left(g^{-1}(\mathbf{y})\right)\left|\det J_{g^{-1}}(\mathbf{y})\right|,\qquad J_{g^{-1}}=\left[\frac{\partial (g^{-1})_i}{\partial y_j}\right].$$

A determinant is a volume scale factor. The Jacobian matrix turns a tiny cube in $\mathbf{y}$-space into a tiny parallelepiped in $\mathbf{x}$-space, and its determinant says how many times bigger the volume became. The absolute value again absorbs an orientation flip. This is precisely the substitution rule from multivariable calculus, and the probability reading of it is the same equation with a different name for the integrand. The calculus of change of variables handles the integral; here we are doing the probability of it, where the integrand is a density that has to stay non-negative and integrate to one. That determinant is also why the linear algebra of volume is unavoidable as soon as there is more than one variable.

The demo below puts the formula on screen. The left panel is a standard normal density. The right panel is the pushforward, computed two ways at once: the analytic formula as a solid curve, and a histogram of five thousand seeded samples pushed through the same map as translucent bars. If the formula is right the bars sit on the curve. Switch the map between an affine stretch, an exponential, and a logistic squash, and watch the density trade width for height. The bars are not evidence that the simulation is lucky; they are evidence that the formula and the sampler agree, which is what makes the formula usable.

Left: the input density f_X. Right: the pushforward f_Y from the formula, with a seeded-sample histogram on top. The strands underneath carry five quantiles of X across to their images under g.

3

Sampling by inverting the CDF

One uniform stream, every distribution

The change-of-variables formula is a test: given a map, check the output density. But it is also a construction. Ask which map g turns a uniform variable into a variable with a chosen distribution F, and the formula answers immediately. Let U be uniform on [0,1], so its density is the constant one on that interval, and set X=F^{-1}(U), where F^{-1} is the quantile function. Then g=F^{-1}, and the formula gives

$$X=F^{-1}(U),\qquad U\sim\mathrm{Unif}(0,1)\quad\Longrightarrow\quad f_X(x)=1\cdot\left|\frac{d}{dx}F(x)\right|=f_X(x).$$

The derivative of the CDF is the density, so the formula checks itself: the inversion produces exactly the target. This is not a trick with symbols. The quantile function is the map that sends a cumulative probability to the value below which that fraction of the distribution sits. Feed it a uniformly random cumulative probability and you get a random value of X. Every uniform draw is a percentile chosen at random, and the quantile function translates that percentile into a sample.

The demo makes the geometry literal. The curve is the CDF, rising from zero to one. A uniform draw picks a height u on the vertical axis. Follow the horizontal line across to the curve, then drop straight down: the x-coordinate where you land is F^{-1}(u). Because u is uniform, heights are chosen evenly, so values under steep parts of the CDF — where the density is large — are hit often, and values under shallow parts are hit rarely. The rug of samples along the bottom thickens exactly where the target density is tallest. That is the pushforward in action, with the uniform as input and the target as output.

The method is universal and it is why reparameterisation works in modern machine learning. A policy gradient estimator needs to differentiate through a sample, and the inverse-CDF view rewrites the sample as a deterministic function of an independent uniform. The randomness is quarantined in u, the parameters enter only through the map F^{-1}, and the gradient flows where it needs to. The same identity underlies the standard normal sampler: a uniform stream plus an inverse error function, or the Box–Muller rotation, which is a two-dimensional change of variables in disguise.

A uniform draw picks a height on the y-axis, the horizontal line meets the CDF F, and the drop lands at F^{-1}(u). The rug and histogram at the bottom accumulate the landing points.

4

Jensen's inequality

A convex curve always lifts the average

A convex function bends upward, so the straight line between any two of its points lies above the curve. Average the two ends and you get a point on that chord, which is above the curve at the midpoint. That is the picture, and it generalises to any distribution: if g is convex and X has a finite mean, then

$$\mathbb{E}\!\left[g(X)\right]\;\ge\;g\!\left(\mathbb{E}[X]\right).$$

Equality holds precisely when X is constant almost surely, or when g is affine on the support of X. The expectation on the left is an average of g over the values the variable actually takes; the number on the right is g evaluated at the average value. Convexity says those two orders of operation cannot be swapped without losing something, and the loss is the curvature times the spread.

The proof is a tangent. At the point $\mu=\mathbb{E}[X]$, draw the tangent line to g. Convexity guarantees the tangent lies below the curve everywhere, so $g(x)\ge g(\mu)+g'(\mu)(x-\mu)$ for every x. Take expectations of both sides. The linear term has expectation $g'(\mu)(\mathbb{E}[X]-\mu)=0$, and the constant passes through untouched, leaving $\mathbb{E}[g(X)]\ge g(\mu)$. Noticing that the tangent is exactly the quantity that averages to $g(\mu)$ is the whole argument. Reversing the inequality gives the concave case for free.

For a uniform variable on [a,b] and g(x)=x^2 the gap has a closed form: $g(\mathbb{E}[X])=((a+b)/2)^2$ while $\mathbb{E}[g(X)]=(a^2+ab+b^2)/3$, so the difference is (b-a)^2/12 — a quarter of the variance, which is exactly the curvature of x^2 times half the variance. Widen the interval and the gap grows quadratically. The demo lets you move the interval and switch between a square and an exponential, with the tangent and the chord drawn in so the gap is a visible length rather than a symbol.

The solid curve is convex g; the dashed line below it is the tangent at $\mathbb{E}[X]$, and the chord above joins the endpoints. The strip at the bottom is the distribution of X.

Jensen is not a curiosity. It is the reason the arithmetic mean dominates the geometric mean, the reason variance is the second-order price of curvature, and the reason a convex loss with a randomised input has higher expected loss than the same loss at the mean input. In optimisation and estimation it decides which direction an error can go, and in information theory it is the inequality behind the non-negativity of relative entropy, because $-\log$ is convex.

5

Where this shows up

A map is always hiding in the model

Calculus

The same substitution, twice

A change of variables in an integral and a change of variables for a density are the same theorem with different names for the integrand. The calculus of it handles the Jacobian mechanically; here the probability of it adds the requirements that the result integrate to one and stay non-negative.

AI / ML

Reparameterised gradients

Policy-gradient and variational methods sample from a distribution whose parameters they need to differentiate. Writing the sample as F^{-1}(u) moves all the randomness into u and leaves a smooth path to the parameters. This is the reparameterisation behind RLHF objectives and the reason a Gaussian policy can be trained by backprop.

Linear algebra

Volume under a linear map

For a linear map the Jacobian is constant and its determinant is the volume scale. Affine Gaussian changes of variables, whitening, and covariance propagation all reduce to $|\det A|$, which is why the linear algebra of determinants is the multivariate half of this formula.

6

Cheat sheet

Every formula in one place

IdeaFormulaReading
Monotone change of variables$f_Y(y)=f_X(g^{-1}(y))\,|(g^{-1})'(y)|$Probability is conserved; the density rescales by the local stretch.
Increasing gf_Y(y)=f_X(x)/g'(x), x=g^{-1}(y)Steep g means thin output density.
Decreasing gf_Y(y)=f_X(x)/|g'(x)|The absolute value handles the reflection.
Multivariate$f_{\mathbf{Y}}(\mathbf{y})=f_{\mathbf{X}}(g^{-1}(\mathbf{y}))\,|\det J_{g^{-1}}(\mathbf{y})|$The Jacobian determinant is the local volume scale.
Inverse-CDF samplingX=F^{-1}(U), $U\sim\mathrm{Unif}(0,1)$One uniform draws a random percentile.
Derivative of the CDF$\frac{d}{dx}F(x)=f_X(x)$Makes the inversion formula check itself.
Jensen, convex$\mathbb{E}[g(X)]\ge g(\mathbb{E}[X])$Curvature pushes the average up; equality only if X is constant.
Jensen, concave$\mathbb{E}[g(X)]\le g(\mathbb{E}[X])$Same statement with the inequality reversed.
Quadratic gap$\mathbb{E}[X^2]-\mathbb{E}[X]^2=\mathrm{Var}(X)$The Jensen gap for x^2 is the variance.
7

Further reading

Where to go deeper

8

Check your understanding

0/6 answered