Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

Byte-pair encoding, in one idea

Start from bytes, merge what repeats

Byte-pair encoding begins with the smallest possible alphabet — bytes — and repeatedly merges the pair of symbols that occurs most often in the training text into a new single symbol. Run that loop a few tens of thousands of times and the vocabulary fills up with the units that actually recur: whole common words, frequent fragments like ##ing, ation or pre, and the space that begins a word as its own marker. A token is one of those learned units. Nothing in the definition says "word"; the algorithm optimises frequency, not linguistics.

The consequence is that a word boundary is invisible to the model. Common words cost one token each; a rare word shatters into a stem plus continuation pieces; a string the tokenizer never saw during training shatters all the way down to bytes. Below is a toy model of that behaviour — it is not any vendor's tokenizer, but it produces the same shapes: frequent short words stay whole, long or rare words split, and single non-alphanumeric characters become tokens of their own.

Tokens per character for five kinds of text. The highlighted bar is the passage currently tokenized below.

💡 Durable idea: a tokenizer is a compression scheme tuned on a corpus. Its output is only meaningful relative to the exact vocabulary and merge list that produced it.
2

A token is not a word, and not a character

Three different units, three different counts

It is tempting to treat "token" as a synonym for "word" and multiply by a price. The three units move independently, and their ratio depends entirely on what the text is made of. English prose is the best case: it lands near four characters per token, or roughly 1.3 tokens per word, because its frequent words and morphemes were merged. Source code is worse, because identifiers, punctuation and indentation are not natural-language statistics. Non-Latin scripts are worse still: a script whose characters are individually rare in the training mix gets tokenized nearly character by character, so a single Chinese or Japanese word can cost several tokens where its English translation costs one.

Numbers are the quiet trap. Whether 128000, 3.14159 or 2026-09-15 is one token or four depends on what the merge list happened to capture, and the answer differs between one model family and the next. This is why arithmetic is a poor first task for a language model, and why "the model got the digit wrong" is often a tokenization fact before it is a reasoning fact.

English prose
≈ 4 characters / token

≈ 1.3 tokens / word. Frequent words and morphemes merged whole.

Source code & URLs
≈ 2–3 characters / token

Identifiers, punctuation and query strings fragment into many pieces.

Non-Latin scripts, emoji, digits
≈ 1 character / token or worse

Rare symbols fall back toward bytes; a number's split is arbitrary.

⚠️ Trap: estimating cost, or a context budget, from a character count silently inflates or deflates by a factor of four depending on the language and the payload. Count tokens, not letters.
3

Tokenizers differ per model family

Counts and ids are not portable

Each model family trains its own tokenizer over its own data mix, so it has a different vocabulary size, a different merge list, and therefore a different number of tokens for the same sentence. The integer id that stands for a token is even less meaningful: it is a serial number into that family's embedding table, assigned by construction order, not by any property of the text. Token #18447 in one family and #18447 in another have nothing whatsoever to do with each other.

The demo below sends the current passage through two mock tokenizers with different hash functions. The token strings are identical; the ids are entirely different. That is the whole point — a count from a tokenizer you are not actually going to call is a guess, and the estimate is only trustworthy when it comes from the same model version you will run.

Two mock vocabularies number the same tokens differently. IDs are arbitrary serial numbers into a family's embedding table.

Token counts are a property of a tokenizer, not of text. The passage in the bench above is re-tokenized here, so editing it changes both demos.

💡 Practical rule: count tokens with the exact tokenizer for the exact model version you will call, and re-count after every model upgrade. A token-count regression test is cheap insurance against a silent context overflow.
4

The context window is a budget

Line items, and a reserve for the answer

The context window is the maximum number of tokens the model can attend to in a single call — the prompt plus the completion. Everything the model knows about this request must fit inside it: the system instructions, the few-shot examples, the retrieved passages, the conversation history, and then the answer it is about to write. Treat it as a budget with named line items, and reserve room for the output first, because a truncated answer is usually worse than a shorter prompt.

Drag the line items below. The stacked column is one prompt's composition against its window, the dashed line is the ceiling, and the readout warns before the input plus the reserve overruns it.

One stacked column: system, examples, retrieved, history, the reserved output, and what is left unused.

Nominal window sizes are a marketing number, not a working capacity. The middle of a long context is where information gets lost: accuracy is highest for material at the very start and the very end and sags in between, so a fact parked at position 40,000 of a 128,000-token window can be effectively absent. . Read the window as an upper bound you should rarely approach.

⚠️ Trap: "the window got bigger" is not the same as "you can use it all." A 1M-token window does not give you a 1M-token working memory; it gives you a larger budget inside which the same primacy-and-recency curve still applies.

Two more line items deserve names. Prefix caching lets a provider reuse the computation for a byte-identical prompt prefix, so keeping stable content first and volatile content last is a budget decision as well as a latency one.

5

What fits, given a cost per item

Turn the budget into a count

Most context decisions are really a packing problem: given a window, a reserve for the answer, some fixed overhead, and a per-item token cost, how many items fit? Each retrieved chunk, few-shot example or tool result has a price in tokens, and the number that fits falls as that price rises. The demo sweeps the per-item cost and plots the count that fits; the marker is the cost you currently have selected.

The design lesson is to measure items in tokens before you decide how many to buy. Ten chunks of 800 tokens and twenty chunks of 400 tokens occupy the same budget but behave very differently in retrieval quality — and neither is free.

Items that fit versus the token cost of one item. The dashed line marks your current choice.

💡 Carry this forward: every later part spends this budget — examples, retrieved chunks, tool definitions, and the model's own reasoning. The lever is almost never the input; as the price table below shows, it is the output.

Cheat sheet

QuestionShort answer
What is a token?A learned subword unit. BPE starts from bytes and merges the most frequent pairs repeatedly.
Is a token a word?No. Common words are one token; rare or foreign words split into several; scripts and numbers vary.
How big is an English token?Roughly four characters, or about 1.3 tokens per word — for English prose only.
Can I reuse a token count across models?No. Tokenizers differ per family; count with the exact model version you will call.
What does a token id mean?Nothing. It is an arbitrary serial number into that family's embedding table.
What is the context window?The most tokens one call can attend to — prompt plus completion. A budget with line items.
What should I reserve first?The output. A truncated answer is usually worse than a shorter prompt.
Does a bigger window mean more usable memory?No. Accuracy sags in the middle of a long context; the window is an upper bound you rarely want.
How do I decide how many chunks to send?Convert each item to tokens and solve the packing problem against the usable budget.

Further reading

6

Check your understanding

0/4 answered