Agents in Action, Interactively
How to design, evaluate and ship an agentic application - workflows and the loop, tool design, planning, memory and compaction, the multi-agent argument, evaluation from error analysis to judges, fine-tuning for products, and production hardening against injection, cost and drift.
The Building with LLMs volume ends with an agent that can search and read a documentation corpus. This volume asks the harder question: whether that agent should exist at all, how to make it reliable, and how to know. It starts at the workflow-versus-agent decision, moves through tool design, planning, memory and the multi-agent argument, then spends three acts on the part that is actually hard — evaluation — before tuning and shipping.
The throughline is the same documentation assistant, grown from a single tool call into an evaluated, hardened deployment. Where an idea belongs to the training or serving layer instead, this volume links out rather than re-teaching: Part 14 of the training guide owns tool use as a post-training problem, and Part 18 of the serving guide owns agent traffic as a workload shape. Keep the glossary and the decision tables with it open.
Already know some of this? Start here
| If you... | Start at |
|---|---|
| Are deciding whether you need an agent at all | Part 1 — Workflows before agents |
| Your agent calls tools but picks the wrong one | Part 2 — Tool use: the loop, and tool design as the craft |
| Your context fills up before the task finishes | Part 5 — Memory and compaction |
| Are arguing about whether to add a second agent | Part 6 — Multi-agent: the argument |
| Have traces but no idea what is failing | Part 8 — Why eval is the hard part |
| Want to judge outputs at scale without fooling yourself | Part 10 — LLM-as-judge and its biases |
| Are about to put a model behind an untrusted input | Part 15 — Prompt injection and the lethal trifecta |
The parts
The five workflow patterns, the complexity/value/viability/cost-of-error gate, and why most 'agents' should be workflows.
Schemas, tool_use/tool_result, parallel calls, and errors as data; few non-overlapping tools, token-efficient returns, and error messages written for the model.
ReAct, Plan-and-Execute, Reflexion, ToT and LATS — what survived to production, and why the ideas outlived the algorithms.
The protocol and its stateless core; tool-definition bloat; SKILL.md's three loading stages; and why the context window is not a data bus.
Working, episodic, semantic and procedural memory; compaction versus context editing versus external notes; MemGPT/Letta, Mem0 and temporal graphs; compaction loss is adversarial.
Orchestrator-worker against single-thread; Anthropic's research system against Cognition's objection, and the read-heavy-and-independent versus shared-mutable-state reconciliation.
Context folding, step budgets, loop detection, sandboxing, permission gates, computer use, deep research as the composite case, and the harness as a source of variance.
Look at your data; open and axial coding; the failure taxonomy and its Pareto chart; build your own trace viewer.
Golden sets, coverage matrices, held-out test sets, code checks before judges, regression gating and significance.
Pointwise versus pairwise, rubrics from the taxonomy, position/verbosity/self-preference bias and its mitigation, validating the judge against humans, and IAA first.
Component-wise metrics, faithfulness and context precision/recall; end-state versus trajectory, pass^k, cost per completed task; SWE-bench Verified, τ²-bench, GAIA and OSWorld — and why public numbers don't transfer.
OTel GenAI semantics, logging retrieved context and prompt versions, A/B power, implicit signals, and cost and latency on the same dashboard.
The decision tree from prompt to RAG to eval to tune to RL; knowledge versus behaviour; the upgrade treadmill; your asset is the dataset, not the weights.
Curation, synthetic data and its diversity collapse, rejection sampling, LoRA/QLoRA operations, DPO/KTO/ORPO, RLVR and GRPO, reward hacking, and adapter serving.
Direct versus indirect injection; private data, untrusted content and an exfiltration channel; why no system prompt fixes it.
The six design patterns, capability-based information flow, dual-LLM; rails and classifiers as defence in depth; PII redaction before logging; GCG, PAIR and AgentDojo in CI.
Timeouts, backoff, circuit breakers and degraded modes; the ordered cost-lever list; three cache layers and semantic caching's correctness cost; routing and when not to cascade; streaming and speculative UX.
Prompts as versioned artifacts, canaries and rollback, risk-tiered autonomy, UX for uncertainty, and capturing user edits as signal.
Multimodal and voice and the ~800 ms budget; the small-model/on-device shift; A2A/AP2/x402 and what protocols cannot express; reliability compounding; how to keep up without chasing.