Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

The Building with LLMs volume ends with an agent that can search and read a documentation corpus. This volume asks the harder question: whether that agent should exist at all, how to make it reliable, and how to know. It starts at the workflow-versus-agent decision, moves through tool design, planning, memory and the multi-agent argument, then spends three acts on the part that is actually hard — evaluation — before tuning and shipping.

The throughline is the same documentation assistant, grown from a single tool call into an evaluated, hardened deployment. Where an idea belongs to the training or serving layer instead, this volume links out rather than re-teaching: Part 14 of the training guide owns tool use as a post-training problem, and Part 18 of the serving guide owns agent traffic as a workload shape. Keep the glossary and the decision tables with it open.

Already know some of this? Start here

If you...Start at
Are deciding whether you need an agent at allPart 1 — Workflows before agents
Your agent calls tools but picks the wrong onePart 2 — Tool use: the loop, and tool design as the craft
Your context fills up before the task finishesPart 5 — Memory and compaction
Are arguing about whether to add a second agentPart 6 — Multi-agent: the argument
Have traces but no idea what is failingPart 8 — Why eval is the hard part
Want to judge outputs at scale without fooling yourselfPart 10 — LLM-as-judge and its biases
Are about to put a model behind an untrusted inputPart 15 — Prompt injection and the lethal trifecta

The parts

Part 1
Workflows before agents

The five workflow patterns, the complexity/value/viability/cost-of-error gate, and why most 'agents' should be workflows.

Part 2
Tool use: the loop, and tool design as the craft

Schemas, tool_use/tool_result, parallel calls, and errors as data; few non-overlapping tools, token-efficient returns, and error messages written for the model.

Part 3
ReAct and the planning family

ReAct, Plan-and-Execute, Reflexion, ToT and LATS — what survived to production, and why the ideas outlived the algorithms.

Part 4
MCP, skills, and code execution

The protocol and its stateless core; tool-definition bloat; SKILL.md's three loading stages; and why the context window is not a data bus.

Part 5
Memory and compaction

Working, episodic, semantic and procedural memory; compaction versus context editing versus external notes; MemGPT/Letta, Mem0 and temporal graphs; compaction loss is adversarial.

Part 6
Multi-agent: the argument

Orchestrator-worker against single-thread; Anthropic's research system against Cognition's objection, and the read-heavy-and-independent versus shared-mutable-state reconciliation.

Part 7
Long-horizon agents and harness design

Context folding, step budgets, loop detection, sandboxing, permission gates, computer use, deep research as the composite case, and the harness as a source of variance.

Part 8
Why eval is the hard part, and error analysis

Look at your data; open and axial coding; the failure taxonomy and its Pareto chart; build your own trace viewer.

Part 9
Datasets, deterministic checks and CI

Golden sets, coverage matrices, held-out test sets, code checks before judges, regression gating and significance.

Part 10
LLM-as-judge and its biases

Pointwise versus pairwise, rubrics from the taxonomy, position/verbosity/self-preference bias and its mitigation, validating the judge against humans, and IAA first.

Part 11
Evaluating RAG and agents

Component-wise metrics, faithfulness and context precision/recall; end-state versus trajectory, pass^k, cost per completed task; SWE-bench Verified, τ²-bench, GAIA and OSWorld — and why public numbers don't transfer.

Part 12
Observability, online eval, and evals as the moat

OTel GenAI semantics, logging retrieved context and prompt versions, A/B power, implicit signals, and cost and latency on the same dashboard.

Part 13
When not to fine-tune

The decision tree from prompt to RAG to eval to tune to RL; knowledge versus behaviour; the upgrade treadmill; your asset is the dataset, not the weights.

Part 14
Data, distillation, and tuning in practice

Curation, synthetic data and its diversity collapse, rejection sampling, LoRA/QLoRA operations, DPO/KTO/ORPO, RLVR and GRPO, reward hacking, and adapter serving.

Part 15
Prompt injection and the lethal trifecta

Direct versus indirect injection; private data, untrusted content and an exfiltration channel; why no system prompt fixes it.

Part 16
Architectural defenses, guardrails and red-teaming

The six design patterns, capability-based information flow, dual-LLM; rails and classifiers as defence in depth; PII redaction before logging; GCG, PAIR and AgentDojo in CI.

Part 17
Reliability, cost and latency engineering

Timeouts, backoff, circuit breakers and degraded modes; the ordered cost-lever list; three cache layers and semantic caching's correctness cost; routing and when not to cascade; streaming and speculative UX.

Part 18
Shipping probabilistic products

Prompts as versioned artifacts, canaries and rollback, risk-tiered autonomy, UX for uncertainty, and capturing user edits as signal.

Part 19
The frontier, and what's unsolved

Multimodal and voice and the ~800 ms budget; the small-model/on-device shift; A2A/AP2/x402 and what protocols cannot express; reliability compounding; how to keep up without chasing.

Reference

Start at Part 1 →