Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

A real byte-level BPE tokenizer

Real vocabulary, not a hash

Byte-level BPE (used by GPT-2, GPT-4, and OLMo's dolma2 tokenizer) starts from raw UTF-8 bytes, not characters — so it can tokenize literally any input, including emoji, without ever needing an "unknown token" fallback for common bytes. Below is a vocabulary we trained ourselves: 1,500 merges learned greedily (most frequent adjacent byte-pair first) from ~1.1M characters of Shakespeare's complete works. Token IDs shown are the real IDs assigned during training — the order each token was learned — not a hash.

⚠️ Try lowering the merge slider to 0: every token becomes a single byte and the vocabulary shrinks to ~65 entries (only the byte values that appear anywhere in Shakespeare — pure ASCII). Raise it back up and common words collapse into single tokens. This is the exact trade-off real tokenizer vocab-size choices make: fewer merges → longer sequences, smaller embedding table; more merges → shorter sequences, bigger embedding table (more rows to learn).
2

Why token cost varies by language

The pricing claim, made visible

Our vocabulary above only ever saw English (and only ASCII bytes). Watch what happens when the same short sentence — translated — hits this tokenizer: languages and scripts it never trained on fragment into far more tokens, because it has no merges for them and, for bytes outside ASCII entirely, not even a base vocabulary entry (it falls back to one token per raw byte). This is a real, structural reason non-English and non-Latin-script text costs more tokens — and hence more money — against API pricing, not a made-up claim.

💡 Production tokenizers (like OLMo's dolma2 or GPT-4's cl100k) reserve one base-vocabulary slot for every possible byte value up front (256 of them) specifically so they never fail outright on unseen scripts — they just fall back to inefficient one-byte-per-token encoding, exactly like ours does above for non-ASCII input.
3

Numbers and whitespace: why arithmetic is hard

A structural quirk

BPE merges whatever byte pairs are most frequent — it has no notion of place value. A 4-digit number might become one token, a 5-digit number three unevenly-split tokens, with no relationship between a number's token boundaries and its digits' actual positions. That misalignment is a real, structural reason LLMs are bad at multi-digit arithmetic unless specifically trained to compensate (e.g. digit-by-digit spacing, or dedicated numeric tokenization).

4

Vocabulary size: a real trade-off

Sequence length vs. embedding table

A bigger vocabulary means shorter token sequences for the same text (good: cheaper attention, which scales quadratically with sequence length) but a bigger embedding table (bad: more parameters spent on the input/output layer instead of the reasoning layers in between). Drag the merge slider from Section 1 down, and watch both numbers below move using our actual vocabulary at each size — not a hypothetical curve.

SchemeUnitUsed by
Byte-level BPEBytes merged by frequencyGPT-2/3/4, OLMo (dolma2)
WordPieceCharacters merged to maximize training-data likelihood, not raw frequencyBERT
Unigram (SentencePiece)Start from a large candidate vocab, prune to maximize corpus likelihoodT5, Llama (SentencePiece mode), ALBERT
5

Embeddings we actually learned

Not hand-placed

Below are real embeddings for the 396 most frequent words in the Shakespeare corpus: we built a word-context co-occurrence matrix (a window of 4 words either side), converted it to pointwise mutual information, reduced it to 12 dimensions with truncated SVD (power iteration, done in plain Python — no ML library), then projected to 2D with PCA purely for this picture. Every position below is a byproduct of real corpus statistics. Click a word to see its actual nearest neighbors (cosine similarity, computed in the full 12-D space).

Click any point. Try "king", "love", "thou", or "death".

Click a word to begin.

These embeddings are genuinely useful, but also genuinely different from a trained transformer's: they're static — one fixed vector per word regardless of context — while a transformer's internal representation of a token is contextual, reshaped at every layer by the other tokens around it (that's exactly what the attention mechanism in Part 3 does). "Bank" near "river" and "bank" near "loan" get the same vector here; they don't in a real model past the first layer.

Vector arithmetic

− +
✓

Cheat sheet

Recap

ConceptWhat it isWhere it fits
Byte-level BPEMerge the most frequent adjacent byte pair, repeatGuarantees every input is tokenizable, even unseen scripts
Vocab size trade-offBigger vocab → shorter sequences, bigger embedding tableOLMo 2 uses ~100K tokens
Static embeddingOne learned vector per token IDThe model's input representation, before any attention runs
Contextual representationA token's vector after passing through transformer layersReshaped per-token by attention over the other tokens present (Part 3)
📚

Further reading

References

?

Check your understanding

0/5 answered
With tokens and embeddings as raw material, the next question is the function that turns a sequence of them into predictions: the transformer block. Continue: the transformer, block by block →