Tokenization & Embeddings
Part 1 used a small illustrative tokenizer and a hand-placed 2D embedding picture, both flagged as such. This part replaces both with the real thing: a byte-level BPE vocabulary we actually trained (1,500 merges, on public-domain text — same algorithm GPT-2 and OLMo use), and word vectors we actually learned from co-occurrence statistics and reduced with real linear algebra, not hand-tuned coordinates.
A real byte-level BPE tokenizer
Real vocabulary, not a hash
Byte-level BPE (used by GPT-2, GPT-4, and OLMo's dolma2 tokenizer) starts from raw UTF-8 bytes, not characters — so it can tokenize literally any input, including emoji, without ever needing an "unknown token" fallback for common bytes. Below is a vocabulary we trained ourselves: 1,500 merges learned greedily (most frequent adjacent byte-pair first) from ~1.1M characters of Shakespeare's complete works. Token IDs shown are the real IDs assigned during training — the order each token was learned — not a hash.
Why token cost varies by language
The pricing claim, made visible
Our vocabulary above only ever saw English (and only ASCII bytes). Watch what happens when the same short sentence — translated — hits this tokenizer: languages and scripts it never trained on fragment into far more tokens, because it has no merges for them and, for bytes outside ASCII entirely, not even a base vocabulary entry (it falls back to one token per raw byte). This is a real, structural reason non-English and non-Latin-script text costs more tokens — and hence more money — against API pricing, not a made-up claim.
Numbers and whitespace: why arithmetic is hard
A structural quirk
BPE merges whatever byte pairs are most frequent — it has no notion of place value. A 4-digit number might become one token, a 5-digit number three unevenly-split tokens, with no relationship between a number's token boundaries and its digits' actual positions. That misalignment is a real, structural reason LLMs are bad at multi-digit arithmetic unless specifically trained to compensate (e.g. digit-by-digit spacing, or dedicated numeric tokenization).
Vocabulary size: a real trade-off
Sequence length vs. embedding table
A bigger vocabulary means shorter token sequences for the same text (good: cheaper attention, which scales quadratically with sequence length) but a bigger embedding table (bad: more parameters spent on the input/output layer instead of the reasoning layers in between). Drag the merge slider from Section 1 down, and watch both numbers below move using our actual vocabulary at each size — not a hypothetical curve.
| Scheme | Unit | Used by |
|---|---|---|
| Byte-level BPE | Bytes merged by frequency | GPT-2/3/4, OLMo (dolma2) |
| WordPiece | Characters merged to maximize training-data likelihood, not raw frequency | BERT |
| Unigram (SentencePiece) | Start from a large candidate vocab, prune to maximize corpus likelihood | T5, Llama (SentencePiece mode), ALBERT |
Embeddings we actually learned
Not hand-placed
Below are real embeddings for the 396 most frequent words in the Shakespeare corpus: we built a word-context co-occurrence matrix (a window of 4 words either side), converted it to pointwise mutual information, reduced it to 12 dimensions with truncated SVD (power iteration, done in plain Python — no ML library), then projected to 2D with PCA purely for this picture. Every position below is a byproduct of real corpus statistics. Click a word to see its actual nearest neighbors (cosine similarity, computed in the full 12-D space).
Click any point. Try "king", "love", "thou", or "death".
These embeddings are genuinely useful, but also genuinely different from a trained transformer's: they're static — one fixed vector per word regardless of context — while a transformer's internal representation of a token is contextual, reshaped at every layer by the other tokens around it (that's exactly what the attention mechanism in Part 3 does). "Bank" near "river" and "bank" near "loan" get the same vector here; they don't in a real model past the first layer.
Vector arithmetic
Cheat sheet
Recap
| Concept | What it is | Where it fits |
|---|---|---|
| Byte-level BPE | Merge the most frequent adjacent byte pair, repeat | Guarantees every input is tokenizable, even unseen scripts |
| Vocab size trade-off | Bigger vocab → shorter sequences, bigger embedding table | OLMo 2 uses ~100K tokens |
| Static embedding | One learned vector per token ID | The model's input representation, before any attention runs |
| Contextual representation | A token's vector after passing through transformer layers | Reshaped per-token by attention over the other tokens present (Part 3) |
Further reading
References
- Sennrich, Haddow & Birch, "Neural Machine Translation of Rare Words with Subword Units" (2015) — BPE for NLP.
- Radford et al. (OpenAI), GPT-2 report — byte-level BPE.
- Kudo & Richardson, "SentencePiece" (2018).
- Mikolov et al., "Efficient Estimation of Word Representations in Vector Space" (2013) — word2vec, the analogy-arithmetic result this section riffs on.