Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

The question

Uncertainty, priced in bits

Start with a single question: how much does one observation tell you? The honest answer depends entirely on what you expected. "The coin came up heads" is worth one bit if the coin was fair and worth nothing at all if it was double-headed; the same words carry different news because surprise is about the gap between what happened and what you thought would happen. That gap has a natural currency. A message that rules out half the remaining possibilities deserves one bit, one that rules out three quarters deserves about 0.42 bits, and one that was already certain deserves zero. The function that obeys this accounting is the negative logarithm, so we define the information content of an outcome with probability p as $-\log p$.

Uncertainty is then the average of that surprise. Weight each outcome's information content by how often it occurs and add up: that weighted average is the entropy of the distribution. A fair coin sits at one bit, a fair eight-sided die at three bits, a deterministic variable at zero. Entropy is largest when every outcome is equally likely, and it is exactly what a good compression scheme needs per symbol — not as an analogy but as a theorem. Claude Shannon's source-coding theorem says that no lossless code can use fewer than H bits per symbol on average, and that codes exist which come arbitrarily close.

Once you accept that a distribution has a price tag, three questions follow immediately. What does it cost to describe data using the wrong distribution? How far apart are two distributions? And how much of one variable is contained in another? The answers are cross-entropy, the KL divergence, and mutual information, and they are tightly braided: cross-entropy is entropy plus KL, and mutual information is a KL divergence between a joint distribution and the product of its own marginals. This part builds all four, with a picture for each.

The pictures matter because the formulas hide their own asymmetry. KL divergence is not a distance: swapping its arguments gives a different number and a different behaviour. One direction refuses to put mass where the truth has none, and produces fits that cover every mode; the other refuses to put mass where the truth is empty, and produces fits that commit to a single mode. That single distinction is why variational inference under-fits variance, why language-model training does not, and why reinforcement learning keeps a KL leash on its policy. So we will draw both fits and let you see the difference rather than recite it.

💡 By the end of this part you'll see why entropy is the average surprise in bits, why cross-entropy is what a wrong model costs, why KL divergence is that excess and why its two directions behave oppositely, and how mutual information packages the same logarithm into a measure of dependence.
2

Entropy: the average surprise

Shape a distribution and watch H fall

Let $p=(p_1,\dots,p_K)$ be a probability distribution over K symbols. Its entropy is the expected information content of one draw,

$$H(p)=-\sum_{i=1}^{K} p_i\log p_i,\qquad 0\le H(p)\le \log K,$$

with the convention that $p_i\log p_i=0$ when p_i=0, since $p\log p\to 0$. The base of the logarithm chooses the unit and nothing else: base 2 gives bits, base e gives nats, and the two differ by the constant factor $\ln 2$. We will use bits because a bit is a yes/no question you can actually ask the world.

Concavity of the logarithm gives the two landmarks. Entropy is largest when the distribution is uniform, where every symbol is equally surprising and $H=\log K$; and it is smallest, at zero, when one symbol has all the mass, because then the outcome is known in advance and no observation carries information. Everything interesting lives between those two poles, and the way a distribution slides from one to the other is exactly how much structure it holds.

The demo below makes that slide tangible. Five bars start equal, so H sits at its ceiling of $\log_2 5\approx 2.32$ bits. Pull any bar up and the entropy drops; the distribution has committed to a symbol, and commitment is the opposite of uncertainty. Push one bar to one and the rest to zero and H collapses to zero, because the message has become a formality. The dashed line marks the uniform level 1/K, the pink outline is a fixed reference distribution q we will use in the next section, and the number above each bar is the optimal integer codelength $\lceil-\log_2 p_i\rceil$ for that symbol.

Blue bars are the distribution p you are shaping; the dashed line is the uniform level 1/K; the pink outline is the reference q used for the cross-entropy readout. The label on each bar is its optimal codelength $\lceil-\log_2 p\rceil$.

p, your distribution q, reference

Watch the two extremes as you drag. At the uniform setting the bars are flat and H equals $\log K$; that is the maximum the formula allows, and no distribution over five symbols can beat it. As you concentrate mass the codelengths above the tall bars shrink and the ones above the short bars grow, because a likely symbol deserves a short code and an unlikely symbol can afford a long one. The average of those codelengths, weighted by p, is what converges to the entropy when the blocks are chosen cleverly enough — which is the content of the next section. For now, remember the shape: entropy is a measure of how flat the distribution is.

3

Cross-entropy and the codelength reading

What a wrong model costs

Entropy tells you the price of describing symbols that genuinely come from p. Cross-entropy tells you the price when the symbols still come from p but your code was built for a different distribution q. The definition is the same average of logarithms with the wrong distribution inside,

$$H(p,q)=-\sum_{i=1}^{K} p_i\log q_i,\qquad H(p,q)=H(p)+\mathrm{KL}(p\|q)\ge H(p).$$

The codelength reading is the cleanest way to feel it. A reasonable code assigns symbol i a length of about $-\log_2 q_i$ bits, so the average length it achieves on data from p is exactly the cross-entropy. If q=p the average matches the entropy and the code is optimal. If q is wrong the average is larger, and the extra bits per symbol are precisely $\mathrm{KL}(p\|q)$. Being wrong is never free and never in your favour: cross-entropy can only exceed entropy, with equality only when the two distributions agree.

This is why cross-entropy is the loss function of choice for classifiers and language models. When a model outputs a distribution q over the next token and the data supply the true next token, minimising $-\log q(\text{token})$ is minimising the average cross-entropy against the empirical distribution. Maximum likelihood and minimum cross-entropy are the same instruction wearing different words, and the gradient it produces is the engine behind the training loops described in the language models part. Reported per token, it is what the evaluation part turns into perplexity by exponentiating it.

Two bookkeeping habits are worth forming. First, units: entropy in nats uses the natural logarithm, and the same quantity in bits is the nat value divided by $\ln 2$. Papers mix them freely, so always check. Second, the readout in the demo is per symbol; for a sequence of n independent symbols the total cost is n times larger, and the law of large numbers says the realised quantity $-\tfrac1n\log p(\text{data})$ concentrates on the entropy. That concentration is the bridge from the abstract formula to a number you can measure from a file, and it is why entropy is a property of a source rather than of a single message.

4

KL divergence and its asymmetry

Two directions, two behaviours

Subtract the entropy from the cross-entropy and the excess bits per symbol get a name: the Kullback–Leibler divergence. It measures how much q has to lose relative to the truth p, and it is non-negative by Jensen's inequality on the concave logarithm.

$$\mathrm{KL}(p\|q)=\sum_{i=1}^{K} p_i\log\frac{p_i}{q_i}=H(p,q)-H(p)\ge 0,\qquad \mathrm{KL}(p\|q)=0\iff p=q.$$

It is not a distance. It is asymmetric, and it violates the triangle inequality, so $\mathrm{KL}(p\|q)$ and $\mathrm{KL}(q\|p)$ are genuinely different numbers with genuinely different preferences. The reason is which distribution does the weighting. The forward divergence $\mathrm{KL}(p\|q)$ weights the log-ratio by p, so it is blind to regions where p is zero — but it punishes q hard for being near zero anywhere p has mass, because $\log(p/q)$ blows up. Minimising it therefore spreads q to cover all of p; it is mode-covering and zero-avoiding. The reverse divergence $\mathrm{KL}(q\|p)$ weights by q, so it is blind to regions where q is zero — but it punishes q for putting mass where p is near zero. Minimising it makes q retreat to a well-supported region and ignore the rest; it is mode-seeking and zero-forcing.

The demo pins this down on the hardest case for a single Gaussian: a target with two separated modes. A single bell curve cannot match a two-humped distribution exactly, so the choice of direction changes which compromise it picks. The reverse fit locks onto one mode and reports a sharp, confident distribution that ignores the other hump entirely. The forward fit stretches to cover both, placing most of its mass in the valley between them where the target is nearly empty. Neither is "right"; they answer different questions about what a fit is for.

The blue curve is the bimodal target p. The pink dashed curve minimises $\mathrm{KL}(q\|p)$ and commits to one mode; the gold dashed curve minimises $\mathrm{KL}(p\|q)$ and straddles both.

target p reverse KL, mode-seeking forward KL, mode-covering

The consequences reach well past this toy. Variational inference, the workhorse behind latent-variable models and the ELBO of the latent variables part, minimises $\mathrm{KL}(q\|p)$ over a tractable family q, so its posteriors are systematically under-dispersed and lock onto a single explanation. Maximum likelihood, and therefore most deep learning, minimises $\mathrm{KL}(p_{\text{data}}\|q)$, so its models spread to cover the data at the risk of putting mass in implausible places. And when a language model is fine-tuned against a reward, the RLHF part adds a KL penalty holding the policy near its reference, which is exactly a forward divergence used as a leash: too far and the model forgets its language; too close and it cannot improve.

5

Mutual information

What one variable knows about another

Everything so far has been about a single distribution. Now take two variables and ask how much one tells you about the other, and the same logarithm answers immediately: compare the joint distribution p(x,y) with the product of the marginals p(x)p(y), which is what the joint would be if the variables had nothing to do with each other. The discrepancy is a KL divergence, and it is called the mutual information.

$$I(X;Y)=\sum_{x,y}p(x,y)\log\frac{p(x,y)}{p(x)\,p(y)}=\mathrm{KL}\big(p(x,y)\,\|\,p(x)p(y)\big)=H(X)+H(Y)-H(X,Y).$$

Three properties fall straight out of that definition. It is symmetric, because the log-ratio and the weighting are symmetric in x and y. It is non-negative, being a KL divergence, and it is zero exactly when the joint factorises, which is the definition of independence. And it can be read as the expected log-likelihood ratio between "these two variables are related" and "these two variables are unrelated" — the evidence, per observation, that the relationship is real. Subtracting both sides of the chain rule $H(X,Y)=H(X)+H(X\mid Y)$ gives the equivalent reading $I(X;Y)=H(X)-H(X\mid Y)$: how many bits of X the observation of Y removes.

The demo builds a joint distribution on a grid and lets you dial its correlation. When the slider sits at zero the two variables are independent by construction, the joint is exactly the outer product of its marginals, the second heatmap coincides with the first, and the mutual information falls to zero. Turn the correlation up and the joint concentrates along a diagonal, which is a pattern the marginals cannot see; the gap between the two heatmaps is the mutual information, and it climbs steadily with the strength of the dependence. Drag the slider through zero and watch the quantity stay pinned at the floor rather than going negative: a divergence never does.

Joint distribution p(x,y) over an $8\times 8$ grid, darker cells carrying more mass. The correlation slider rotates the ridge from the anti-diagonal through independence to the main diagonal.

The product of the marginals $p(x)\,p(y)$, drawn on the same colour scale. Wherever it differs from the joint above, the joint is carrying dependence the marginals cannot.

low mass high mass

Mutual information is the quantity behind channel capacity, feature selection, and the attention weights of a transformer, all of which are asking the same question in different clothes: how many bits does this variable hand me about that one? It is also the first genuinely graphical quantity in this volume. Once you can factor a joint distribution, you can read independence and conditional independence off the structure of a model, and the next part turns exactly that factoring into Bayes nets and factor graphs.

6

Where this shows up

The same logarithm, four costumes

AI / ML

Language-model loss

A next-token model is trained by minimising the cross-entropy between its predicted distribution and the observed token. That is average negative log-likelihood, it decomposes over positions by the chain rule, and the language models part derives the exact form the gradient takes once teacher forcing is in play.

AI / ML

Perplexity and calibration

Perplexity is 2^{H}, the effective number of equally likely choices per token; lower is better and always at least the true entropy. When a model's stated probabilities drift from the observed frequencies, cross-entropy and calibration diverge, which is the gap the evaluation part is built to measure.

AI / ML

The KL leash in RLHF

Reward maximisation alone drives a policy off the manifold of fluent text, so the objective adds a KL penalty against a frozen reference model. The RLHF part shows the coefficient as the dial between "say anything that scores" and "stay recognisably the model you started with."

Math

Densities and their Jacobians

For continuous variables the sum becomes an integral and the entropy picks up a reference scale, but every identity survives. Pushing a variable through a map rescales its density by the Jacobian, and the calculus of probability is where that change of variables is worked out.

7

Cheat sheet

Every formula in one place

IdeaFormulaIntuition
Informational content$-\log p(x)$The surprise of one outcome, in bits when the base is 2.
Entropy$H(p)=-\sum_i p_i\log p_i$The average surprise; how flat the distribution is.
Entropy bound$0\le H(p)\le\log K$Uniform is maximum, a point mass is zero.
Cross-entropy$H(p,q)=-\sum_i p_i\log q_i$Cost of coding draws from p with a code built for q.
KL divergence$\mathrm{KL}(p\|q)=\sum_i p_i\log\frac{p_i}{q_i}=H(p,q)-H(p)$The excess bits; non-negative, zero iff p=q, asymmetric.
Forward KL$\mathrm{KL}(p\|q)$Mode-covering, zero-avoiding; the maximum-likelihood direction.
Reverse KL$\mathrm{KL}(q\|p)$Mode-seeking, zero-forcing; the variational-inference direction.
Mutual information$I(X;Y)=\sum p(x,y)\log\frac{p(x,y)}{p(x)p(y)}$KL between the joint and the product of marginals.
MI identities$I=H(X)+H(Y)-H(X,Y)=H(X)-H(X\mid Y)$Bits one variable removes from the other; zero iff independent.
Optimal codelength$\ell(x)=\lceil-\log_2 p(x)\rceil$, average $\ge H$Likely symbols get short codes; Shannon's bound is the floor.
Perplexity2^{H(p)}Effective number of equally likely choices per symbol.
8

Further reading

Where to go deeper

9

Check your understanding

0/6 answered