Building with LLMs, Interactively
The layer between the model and the product - what a token and a context window really are, how to prompt and engineer context, how retrieval actually works, and how to measure any of it before you ship.
The LLM Training volume covers how a model is made; LLM Serving covers how it is run. This volume owns the layer in between, the one most practitioners actually work in: how to turn a raw text function into an application. It starts from zero — what a token is, what an embedding is, why a context window is a budget rather than a bucket — and ends at retrieval architectures and the arguments still being had about them.
The companion volume, Agents in Action, picks up where this one stops: workflows and agents, tool design, memory and compaction, evaluation, fine-tuning for products, and production hardening. The two share one running example — an assistant over a documentation corpus — grown across both volumes, so the embedding space you build here is the one the agent searches in there. Keep the glossary open for notation, and treat Part 2 of the training guide as the deeper treatment of the tokenizer.
Already know some of this? Start here
| If you... | Start at |
|---|---|
| Have never called an API and want the mental model first | Part 1 — What you're actually building on |
| Are fighting a context limit or a surprise bill | Part 2 — Tokens, windows and budgets, then Part 11 — Caching-aware prompt layout |
| Cannot get the model to return parseable JSON | Part 9 — Structured output and constrained decoding |
| Have a prompt that works but nothing to prove it stays working | Part 5 — Your first eval, in twenty lines |
| Are about to add a vector database to a working keyword search | Part 12 — Why retrieve, and lexical search first |
| Need to choose an index for a hundred million vectors | Part 15 — Approximate nearest neighbour |
| Want the honest case for and against RAG in the long-context era | Part 17 — Beyond chunks |
The parts
The model is a stateless text function; every 'memory' is your code resending text, and the context window is the budget that governs it.
Byte-pair encoding, why a token is not a word, per-family tokenizers, and the context window as a contested budget.
Temperature, top-p and top-k; why temperature 0 is not deterministic; logprobs and the poor calibration of post-RLHF models.
The input/output price asymmetry, TTFT versus TPOT, why an agent loop is quadratic in turns, and effort as a dial.
Five examples, a binary check, a pass rate — deliberately early, so every later part is measurable.
System prompts, delimiters, and data/instruction separation — the first security lesson.
Selection, ordering, and format contagion: the model copies your examples' mistakes.
Wei et al. and Kojima et al. as history; self-consistency as the first test-time-compute pattern; why prescriptive CoT now often hurts.
JSON Schema, strict tools, and grammar/FSM token masking; syntax guarantees are not semantic ones.
Write, select, compress, isolate; compaction, JIT retrieval, structured notes; context rot and lost-in-the-middle; long context versus retrieval.
Prefix matching, stable-content-first ordering, breakpoints, silent invalidators, and measuring cache_read_input_tokens.
Freshness, privacy, citations, ACLs and cost — not 'the model doesn't know things'; BM25 as the brutally strong baseline.
What a vector means, cosine versus relevance, query/document asymmetry, Matryoshka truncation, and MTEB's overfitting problem.
Fixed, recursive, semantic, parent-doc and RAPTOR chunking; boundary effects; and contextual retrieval's measured failure reductions.
Flat, IVF, PQ, HNSW, ScaNN and DiskANN; recall, QPS and memory as a four-way tradeoff; and why filtered ANN is genuinely hard.
RRF versus score normalisation, SPLADE, cross-encoders and late interaction; retrieve wide, rerank narrow.
HyDE, multi-query and step-back; Self-RAG and CRAG; GraphRAG's global questions; ColPali; grep versus embeddings for code; and the long-context argument as an argument.