Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

LLM Training, Interactively

Next-token prediction, autoregressive generation, and why OLMo is the running example for this series.

Part 1
What is a language model

Next-token prediction, autoregressive generation, and why OLMo is the running example for this series.

Part 2
Tokenization & embeddings

Real BPE training, multilingual token cost, and learned (not hand-placed) word vectors.

Part 3
The transformer, block by block

Attention, multi-head, RoPE, residual stream, norms, SwiGLU, parameter counting, and MoE.

Part 4
Building the pretraining dataset

Sourcing, filtering, deduplication, contamination, and mixing at trillion-token scale.

Part 5
Pretraining: the training loop

Loss, perplexity, learning-rate schedules, gradient clipping, and mid-training annealing.

Part 6
Scaling laws & training systems

Chinchilla compute-optimal scaling, memory accounting, and distributed-training parallelism.

Part 7
Supervised fine-tuning & PEFT

Chat templates, loss masking, and LoRA/QLoRA for parameter-efficient fine-tuning.

Part 8
Alignment: RLHF & DPO

Reward models, PPO's clipped objective, DPO's implicit reward, and the KL-reward frontier.

Part 9
Reasoning & RL with verifiable rewards

Chain-of-thought, RLVR, GRPO, and test-time compute scaling.

Part 10
Evaluation

How benchmarks are actually scored, LLM-as-judge, arena ratings, and pass@k.

Part 11
Inference, serving & efficiency

Sampling, the KV cache, prefill vs. decode, batching, and quantization.

Part 12
Guardrails, safety & deployment

Refusal training, layered defenses, and staged release.

Part 13
Long-context & context extension

Stretching RoPE, long-context data, best-fit packing, and why extension is its own late stage.

Part 14
Tool use, function calling & agents

OpenAPI tool specs, environment roles, real vs simulated trajectories, and how tool use is evaluated.

Part 15
Post-training at scale

The three-stage recipe, delta learning, upgraded GRPO, domain verifiers, and RL systems.

Glossary
Glossary & numbers to know

Every bolded term in the series, linked back to where it is introduced, plus a quick-reference card.

Open the 15-part guide →

LLM Serving, Interactively

Why generating one token costs a full sweep of the weights through HBM, and why that single fact shapes everything that follows.

Part 1
The decode loop and the memory wall

Why generating one token costs a full sweep of the weights through HBM, and why that single fact shapes everything that follows.

Part 2
The machine: HBM, interconnect and the accelerator zoo

Bandwidth as the headline number, the all-reduce tax, NVL72 as one big GPU, SRAM architectures, and energy per token.

Part 3
TTFT, TPOT and goodput

The four numbers that describe a serving system, why percentiles beat means, and how to write an SLO you can defend.

Part 4
The KV cache: sizing, GQA, MQA and MLA

The bytes-per-token formula run against real model configs, and the lever it becomes for concurrency.

Part 5
PagedAttention: block tables, fragmentation and preemption

The three kinds of waste in contiguous reservation, blocks and block tables, copy-on-write, and recompute versus swap.

Part 6
Continuous batching and chunked prefill

Why iteration-level batching wins, what batch size does to step time, and the ITL spike chunked prefill trades away.

Part 7
The scheduler: queues, fairness and admission control

What the scheduler decides every step, head-of-line blocking, length prediction, starvation, and when to shed load.

Part 8
Prefix caching, radix trees and the KV hierarchy

Reusing the same prefix across requests, RadixAttention, cache-aware routing, and the tier economics of fetch versus recompute.

Part 9
Inside the engine: kernels, CUDA graphs and the life of a request

The end-to-end request path, FlashAttention tiling, paged kernels, launch overhead, and the engine landscape.

Part 10
Quantization and numerics

Formats from FP16 to MXFP4, activation outliers, why the win is bandwidth not FLOPs, and how to measure the damage.

Part 11
Speculative decoding

Verify many, generate one; acceptance rates, draft models, tree attention, and why engines disable speculation under load.

Part 12
Parallelism for inference: TP, PP, EP, DP-attention, CP

Inference parallelism is not training parallelism, why TP stays inside NVLink, and how to search the sharding space.

Part 13
Prefill/decode disaggregation

One pool, two jobs; sizing the pools, moving KV over RDMA, when disaggregation loses, and what DeepSeek actually runs.

Part 14
Serving mixture-of-experts at scale

Active versus total parameters, all-to-all dispatch, hot experts, EPLB, wide expert parallelism, and why rack-scale exists.

Part 15
Long-context serving

Quadratic prefill, linear KV, ring attention, sparse attention and eviction, and the honest cost per request.

Part 16
Structured output and constrained decoding

Regex to FSM to per-state masks, why JSON needs a pushdown automaton, jump-forward decoding, and the constraint tax.

Part 17
LoRA, quotas and noisy neighbours

One base and many adapters, batched LoRA kernels, adapter paging, token buckets, and fair queueing across tenants.

Part 18
Reasoning, agents and multimodal

Workload shape as a first-class serving parameter: thinking-token budgets, agent turns, image tokens, and realtime voice.

Part 19
Kubernetes, autoscaling, cold starts and routing

Gateways, routers and autoscalers; why GPU utilisation is a lie; the cold-start breakdown; and honest scale-to-zero.

Part 20
Benchmarking, cost and capacity planning

Open versus closed loop, batch non-invariance, dashboard signals, incident signatures, and RPS to GPUs to dollars.

Reference
Glossary & numbers to know

Every term in the series linked back to where it is introduced, plus the accelerator, model, quantization, engine and pricing tables.

Open the 20-part guide →

Building with LLMs, Interactively

The model is a stateless text function; every 'memory' is your code resending text, and the context window is the budget that governs it.

Part 1
What you're actually building on

The model is a stateless text function; every 'memory' is your code resending text, and the context window is the budget that governs it.

Part 2
Tokens, windows and budgets

Byte-pair encoding, why a token is not a word, per-family tokenizers, and the context window as a contested budget.

Part 3
Sampling, nondeterminism and confidence

Temperature, top-p and top-k; why temperature 0 is not deterministic; logprobs and the poor calibration of post-RLHF models.

Part 4
Cost and latency from first principles

The input/output price asymmetry, TTFT versus TPOT, why an agent loop is quadratic in turns, and effort as a dial.

Part 5
Your first eval, in twenty lines

Five examples, a binary check, a pass rate — deliberately early, so every later part is measurable.

Part 6
Instructions, roles and the right altitude

System prompts, delimiters, and data/instruction separation — the first security lesson.

Part 7
Few-shot: examples as specification

Selection, ordering, and format contagion: the model copies your examples' mistakes.

Part 8
Reasoning: CoT, self-consistency, and what reasoning models changed

Wei et al. and Kojima et al. as history; self-consistency as the first test-time-compute pattern; why prescriptive CoT now often hurts.

Part 9
Structured output and constrained decoding

JSON Schema, strict tools, and grammar/FSM token masking; syntax guarantees are not semantic ones.

Part 10
Context engineering

Write, select, compress, isolate; compaction, JIT retrieval, structured notes; context rot and lost-in-the-middle; long context versus retrieval.

Part 11
Caching-aware prompt layout

Prefix matching, stable-content-first ordering, breakpoints, silent invalidators, and measuring cache_read_input_tokens.

Part 12
Why retrieve, and lexical search first

Freshness, privacy, citations, ACLs and cost — not 'the model doesn't know things'; BM25 as the brutally strong baseline.

Part 13
Embeddings and vector similarity

What a vector means, cosine versus relevance, query/document asymmetry, Matryoshka truncation, and MTEB's overfitting problem.

Part 14
Chunking, and contextual retrieval

Fixed, recursive, semantic, parent-doc and RAPTOR chunking; boundary effects; and contextual retrieval's measured failure reductions.

Part 15
Approximate nearest neighbour

Flat, IVF, PQ, HNSW, ScaNN and DiskANN; recall, QPS and memory as a four-way tradeoff; and why filtered ANN is genuinely hard.

Part 16
Hybrid search, fusion and reranking

RRF versus score normalisation, SPLADE, cross-encoders and late interaction; retrieve wide, rerank narrow.

Part 17
Beyond chunks: query transforms, graph, multimodal, code — and the “RAG is dead” debate

HyDE, multi-query and step-back; Self-RAG and CRAG; GraphRAG's global questions; ColPali; grep versus embeddings for code; and the long-context argument as an argument.

Reference
Glossary and selection tables

Every term in the volume linked back to the part that introduces it, plus a retrieval-architecture selection table and a RAG failure-mode to metric table.

Open the 17-part guide →

Agents in Action, Interactively

The five workflow patterns, the complexity/value/viability/cost-of-error gate, and why most 'agents' should be workflows.

Part 1
Workflows before agents

The five workflow patterns, the complexity/value/viability/cost-of-error gate, and why most 'agents' should be workflows.

Part 2
Tool use: the loop, and tool design as the craft

Schemas, tool_use/tool_result, parallel calls, and errors as data; few non-overlapping tools, token-efficient returns, and error messages written for the model.

Part 3
ReAct and the planning family

ReAct, Plan-and-Execute, Reflexion, ToT and LATS — what survived to production, and why the ideas outlived the algorithms.

Part 4
MCP, skills, and code execution

The protocol and its stateless core; tool-definition bloat; SKILL.md's three loading stages; and why the context window is not a data bus.

Part 5
Memory and compaction

Working, episodic, semantic and procedural memory; compaction versus context editing versus external notes; MemGPT/Letta, Mem0 and temporal graphs; compaction loss is adversarial.

Part 6
Multi-agent: the argument

Orchestrator-worker against single-thread; Anthropic's research system against Cognition's objection, and the read-heavy-and-independent versus shared-mutable-state reconciliation.

Part 7
Long-horizon agents and harness design

Context folding, step budgets, loop detection, sandboxing, permission gates, computer use, deep research as the composite case, and the harness as a source of variance.

Part 8
Why eval is the hard part, and error analysis

Look at your data; open and axial coding; the failure taxonomy and its Pareto chart; build your own trace viewer.

Part 9
Datasets, deterministic checks and CI

Golden sets, coverage matrices, held-out test sets, code checks before judges, regression gating and significance.

Part 10
LLM-as-judge and its biases

Pointwise versus pairwise, rubrics from the taxonomy, position/verbosity/self-preference bias and its mitigation, validating the judge against humans, and IAA first.

Part 11
Evaluating RAG and agents

Component-wise metrics, faithfulness and context precision/recall; end-state versus trajectory, pass^k, cost per completed task; SWE-bench Verified, τ²-bench, GAIA and OSWorld — and why public numbers don't transfer.

Part 12
Observability, online eval, and evals as the moat

OTel GenAI semantics, logging retrieved context and prompt versions, A/B power, implicit signals, and cost and latency on the same dashboard.

Part 13
When not to fine-tune

The decision tree from prompt to RAG to eval to tune to RL; knowledge versus behaviour; the upgrade treadmill; your asset is the dataset, not the weights.

Part 14
Data, distillation, and tuning in practice

Curation, synthetic data and its diversity collapse, rejection sampling, LoRA/QLoRA operations, DPO/KTO/ORPO, RLVR and GRPO, reward hacking, and adapter serving.

Part 15
Prompt injection and the lethal trifecta

Direct versus indirect injection; private data, untrusted content and an exfiltration channel; why no system prompt fixes it.

Part 16
Architectural defenses, guardrails and red-teaming

The six design patterns, capability-based information flow, dual-LLM; rails and classifiers as defence in depth; PII redaction before logging; GCG, PAIR and AgentDojo in CI.

Part 17
Reliability, cost and latency engineering

Timeouts, backoff, circuit breakers and degraded modes; the ordered cost-lever list; three cache layers and semantic caching's correctness cost; routing and when not to cascade; streaming and speculative UX.

Part 18
Shipping probabilistic products

Prompts as versioned artifacts, canaries and rollback, risk-tiered autonomy, UX for uncertainty, and capturing user edits as signal.

Part 19
The frontier, and what's unsolved

Multimodal and voice and the ~800 ms budget; the small-model/on-device shift; A2A/AP2/x402 and what protocols cannot express; reliability compounding; how to keep up without chasing.

Reference
Glossary and decision tables

Every term in the volume linked back to the part that introduces it, plus the prompt to RAG to tune decision table, an eval-metric-per-failure-mode table, an injection-defense-pattern table, and a dated benchmark reference card.

Open the 19-part guide →

Generative Media, Interactively

Sampling from a distribution you never see: a 2D dataset you draw yourself, and the line between memorising it and modelling it.

Part 1
What a generative model actually is

Sampling from a distribution you never see: a 2D dataset you draw yourself, and the line between memorising it and modelling it.

Part 2
Destroying an image with noise

The forward process on a procedurally-generated image, q(x_t|x_0) in closed form, and what ᾱ, SNR and the schedule choice actually do.

Part 3
Learning to undo one step

ε-, x₀- and v-prediction, and a real tiny denoiser trained in the browser with its loss curve live.

Part 4
Sampling: noise to picture

The reverse loop from noise to sample; ancestral DDPM against deterministic DDIM, and a step-count slider that shows the trade.

Part 5
The score function

The score field as arrows over a density, Langevin dynamics, the SMLD to DDPM equivalence and the VE versus VP SDEs.

Part 6
Flow matching and rectified flow

Curved paths against straight ones, the velocity-field view of an ODE, and why the frontier trainers made flow matching the default.

Part 7
Latent diffusion

The VAE compressor that shrinks an image before it is diffused, the cost of an eight-times downsample, and the pixel versus latent budget.

Part 8
The U-Net backbone

The resolution pyramid and its skips, residual blocks, where self- and cross-attention sit, and interactive parameter counts.

Part 9
The DiT backbone

Patching a latent into tokens, plain transformer blocks with adaLN-Zero conditioning, and U-Net against DiT.

Part 10
Conditioning and classifier-free guidance

Text encoders and cross- against joint attention, CFG as extrapolation, and what guidance scale costs in diversity.

Part 11
Samplers and schedulers

Euler, Heun and multistep solvers against the ancestral sampler, timestep spacing as a design choice, and step budget against error.

Part 12
Few-step generation

Progressive, consistency, adversarial and shortcut distillation, and what collapsing fifty steps into one or four costs in diversity.

Part 13
Steering a frozen model

LoRA and DreamBooth, ControlNet and IP-Adapter, inpainting and SDEdit editing, and where along the trajectory each intervention bites.

Part 14
Video as a 3D signal

The causal 3-D VAE, spatiotemporal patches, 3-D against factorised attention, and the token budget a clip really costs.

Part 15
Coherent and long

Image and keyframe conditioning, temporal cascades, autoregressive rollout against full-sequence diffusion, drift, and the world-model claim.

Part 16
What audio is, numerically

A waveform, its sample rate, the STFT and the mel scale, and the token-rate arithmetic that makes raw waveform generation hopeless.

Part 17
Neural audio codecs

The encoder, residual vector quantiser and decoder that turn sound into a few hundred discrete tokens a second.

Part 18
Text to speech

From alignment and duration models to codec language models and two-stage synthesis, with cloning and streaming first-chunk latency.

Part 19
Speech to text, and answering out loud

CTC, RNN-T and Whisper, why streaming is hard and the two-pass fix, then voice activity, barge-in and the realtime duplex budget.

Part 20
Evaluating and shipping

FID, FVD and CLIPScore and what each misses, preference models and Elo arenas, memorisation, and provenance under the EU AI Act.

Reference
Glossary & numbers to know

Every term linked back to the part that introduces it, plus the model, sampler and formula tables.

Open the 20-part guide →

Multimodal Models, Interactively

Patching an image into a sequence, position embeddings and the class token, and what a vision transformer gives up against a convolution.

Part 1
Pixels into tokens

Patching an image into a sequence, position embeddings and the class token, and what a vision transformer gives up against a convolution.

Part 2
Caption matches picture

InfoNCE on the similarity matrix, the temperature and batch size it needs, and SigLIP's pairwise sigmoid alternative.

Part 3
What a shared space buys

Zero-shot classification, retrieval and prompt templates, and the compositionality, counting and typographic failures the space inherits.

Part 4
The modality zoo

Audio, video and 3D encoders aligned into one space, what a hub modality binds, and what aligned does and does not mean.

Part 5
Bolting an encoder onto a language model

Linear and MLP projectors, the Q-Former and cross-attention bridges, where image tokens sit, and what each option trains.

Part 6
Resolution and the token budget

Fixed patches, AnyRes tiling and native dynamic resolution, the merge that compresses tokens, and the image's prefill cost.

Part 7
Training a vision-language model

The staged recipe, freezing schedules and data mixtures, what breaks along the way, and preference tuning on multimodal pairs.

Part 8
Why VLMs hallucinate

Objects that were never there, attention sinks on image tokens, the language prior overriding pixels, and the benchmarks that measure it.

Part 9
One trunk, many modalities

Dropping the bridge between towers, native multimodal pretraining, modality embedders into one stream, and modality-aware experts.

Part 10
Video in, long context

Frame sampling against a token budget, temporal position encoding, streaming memory, and the KV arithmetic a clip implies.

Part 11
Listening models

Speech as an encoder's output or as native audio tokens in the trunk, speech instruction following, and paralinguistics.

Part 12
Omni models and the real-time loop

Interleaved token streams, the thinker-and-talker split, listening while speaking, barge-in, and the realtime budget.

Part 13
Understanding and generating

One model that reads and makes images, the tokenizer choice that decides whether it works, and three generation orders.

Part 14
Grounding

Detection and segmentation as coordinate tokens, promptable segmentation, referring expressions, and grounding an interface.

Part 15
Vision-language-action

Actions as text tokens, a flow-matching action expert, chunking, and the latent world models that predict what happens next.

Part 16
Serving and evaluating multimodal systems

Encoder disaggregation and caching, the prefill-and-decode inversion an image causes, safety, and the open problems.

Reference
Glossary & numbers to know

Every term linked back to the part that introduces it, plus the architecture-era, token-budget and evaluation tables.

Open the 16-part guide →