Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

The LLM Training volume covers how a model is made; LLM Serving covers how it is run. This volume owns the layer in between, the one most practitioners actually work in: how to turn a raw text function into an application. It starts from zero — what a token is, what an embedding is, why a context window is a budget rather than a bucket — and ends at retrieval architectures and the arguments still being had about them.

The companion volume, Agents in Action, picks up where this one stops: workflows and agents, tool design, memory and compaction, evaluation, fine-tuning for products, and production hardening. The two share one running example — an assistant over a documentation corpus — grown across both volumes, so the embedding space you build here is the one the agent searches in there. Keep the glossary open for notation, and treat Part 2 of the training guide as the deeper treatment of the tokenizer.

Already know some of this? Start here

If you...Start at
Have never called an API and want the mental model firstPart 1 — What you're actually building on
Are fighting a context limit or a surprise billPart 2 — Tokens, windows and budgets, then Part 11 — Caching-aware prompt layout
Cannot get the model to return parseable JSONPart 9 — Structured output and constrained decoding
Have a prompt that works but nothing to prove it stays workingPart 5 — Your first eval, in twenty lines
Are about to add a vector database to a working keyword searchPart 12 — Why retrieve, and lexical search first
Need to choose an index for a hundred million vectorsPart 15 — Approximate nearest neighbour
Want the honest case for and against RAG in the long-context eraPart 17 — Beyond chunks

The parts

Part 1
What you're actually building on

The model is a stateless text function; every 'memory' is your code resending text, and the context window is the budget that governs it.

Part 2
Tokens, windows and budgets

Byte-pair encoding, why a token is not a word, per-family tokenizers, and the context window as a contested budget.

Part 3
Sampling, nondeterminism and confidence

Temperature, top-p and top-k; why temperature 0 is not deterministic; logprobs and the poor calibration of post-RLHF models.

Part 4
Cost and latency from first principles

The input/output price asymmetry, TTFT versus TPOT, why an agent loop is quadratic in turns, and effort as a dial.

Part 5
Your first eval, in twenty lines

Five examples, a binary check, a pass rate — deliberately early, so every later part is measurable.

Part 6
Instructions, roles and the right altitude

System prompts, delimiters, and data/instruction separation — the first security lesson.

Part 7
Few-shot: examples as specification

Selection, ordering, and format contagion: the model copies your examples' mistakes.

Part 8
Reasoning: CoT, self-consistency, and what reasoning models changed

Wei et al. and Kojima et al. as history; self-consistency as the first test-time-compute pattern; why prescriptive CoT now often hurts.

Part 9
Structured output and constrained decoding

JSON Schema, strict tools, and grammar/FSM token masking; syntax guarantees are not semantic ones.

Part 10
Context engineering

Write, select, compress, isolate; compaction, JIT retrieval, structured notes; context rot and lost-in-the-middle; long context versus retrieval.

Part 11
Caching-aware prompt layout

Prefix matching, stable-content-first ordering, breakpoints, silent invalidators, and measuring cache_read_input_tokens.

Part 12
Why retrieve, and lexical search first

Freshness, privacy, citations, ACLs and cost — not 'the model doesn't know things'; BM25 as the brutally strong baseline.

Part 13
Embeddings and vector similarity

What a vector means, cosine versus relevance, query/document asymmetry, Matryoshka truncation, and MTEB's overfitting problem.

Part 14
Chunking, and contextual retrieval

Fixed, recursive, semantic, parent-doc and RAPTOR chunking; boundary effects; and contextual retrieval's measured failure reductions.

Part 15
Approximate nearest neighbour

Flat, IVF, PQ, HNSW, ScaNN and DiskANN; recall, QPS and memory as a four-way tradeoff; and why filtered ANN is genuinely hard.

Part 16
Hybrid search, fusion and reranking

RRF versus score normalisation, SPLADE, cross-encoders and late interaction; retrieve wide, rerank narrow.

Part 17
Beyond chunks: query transforms, graph, multimodal, code — and the “RAG is dead” debate

HyDE, multi-query and step-back; Self-RAG and CRAG; GraphRAG's global questions; ColPali; grep versus embeddings for code; and the long-context argument as an argument.

Reference

Start at Part 1 →