AI
Six interactive guides to how a language model is trained, how one is served at scale, how to build with LLMs, how agents act in the world, and how models are made to see, hear and generate media.
LLM Training, Interactively
Next-token prediction, autoregressive generation, and why OLMo is the running example for this series.
Next-token prediction, autoregressive generation, and why OLMo is the running example for this series.
Real BPE training, multilingual token cost, and learned (not hand-placed) word vectors.
Attention, multi-head, RoPE, residual stream, norms, SwiGLU, parameter counting, and MoE.
Sourcing, filtering, deduplication, contamination, and mixing at trillion-token scale.
Loss, perplexity, learning-rate schedules, gradient clipping, and mid-training annealing.
Chinchilla compute-optimal scaling, memory accounting, and distributed-training parallelism.
Chat templates, loss masking, and LoRA/QLoRA for parameter-efficient fine-tuning.
Reward models, PPO's clipped objective, DPO's implicit reward, and the KL-reward frontier.
Chain-of-thought, RLVR, GRPO, and test-time compute scaling.
How benchmarks are actually scored, LLM-as-judge, arena ratings, and pass@k.
Sampling, the KV cache, prefill vs. decode, batching, and quantization.
Refusal training, layered defenses, and staged release.
Stretching RoPE, long-context data, best-fit packing, and why extension is its own late stage.
OpenAPI tool specs, environment roles, real vs simulated trajectories, and how tool use is evaluated.
The three-stage recipe, delta learning, upgraded GRPO, domain verifiers, and RL systems.
Every bolded term in the series, linked back to where it is introduced, plus a quick-reference card.
LLM Serving, Interactively
Why generating one token costs a full sweep of the weights through HBM, and why that single fact shapes everything that follows.
Why generating one token costs a full sweep of the weights through HBM, and why that single fact shapes everything that follows.
Bandwidth as the headline number, the all-reduce tax, NVL72 as one big GPU, SRAM architectures, and energy per token.
The four numbers that describe a serving system, why percentiles beat means, and how to write an SLO you can defend.
The bytes-per-token formula run against real model configs, and the lever it becomes for concurrency.
The three kinds of waste in contiguous reservation, blocks and block tables, copy-on-write, and recompute versus swap.
Why iteration-level batching wins, what batch size does to step time, and the ITL spike chunked prefill trades away.
What the scheduler decides every step, head-of-line blocking, length prediction, starvation, and when to shed load.
Reusing the same prefix across requests, RadixAttention, cache-aware routing, and the tier economics of fetch versus recompute.
The end-to-end request path, FlashAttention tiling, paged kernels, launch overhead, and the engine landscape.
Formats from FP16 to MXFP4, activation outliers, why the win is bandwidth not FLOPs, and how to measure the damage.
Verify many, generate one; acceptance rates, draft models, tree attention, and why engines disable speculation under load.
Inference parallelism is not training parallelism, why TP stays inside NVLink, and how to search the sharding space.
One pool, two jobs; sizing the pools, moving KV over RDMA, when disaggregation loses, and what DeepSeek actually runs.
Active versus total parameters, all-to-all dispatch, hot experts, EPLB, wide expert parallelism, and why rack-scale exists.
Quadratic prefill, linear KV, ring attention, sparse attention and eviction, and the honest cost per request.
Regex to FSM to per-state masks, why JSON needs a pushdown automaton, jump-forward decoding, and the constraint tax.
One base and many adapters, batched LoRA kernels, adapter paging, token buckets, and fair queueing across tenants.
Workload shape as a first-class serving parameter: thinking-token budgets, agent turns, image tokens, and realtime voice.
Gateways, routers and autoscalers; why GPU utilisation is a lie; the cold-start breakdown; and honest scale-to-zero.
Open versus closed loop, batch non-invariance, dashboard signals, incident signatures, and RPS to GPUs to dollars.
Every term in the series linked back to where it is introduced, plus the accelerator, model, quantization, engine and pricing tables.
Building with LLMs, Interactively
The model is a stateless text function; every 'memory' is your code resending text, and the context window is the budget that governs it.
The model is a stateless text function; every 'memory' is your code resending text, and the context window is the budget that governs it.
Byte-pair encoding, why a token is not a word, per-family tokenizers, and the context window as a contested budget.
Temperature, top-p and top-k; why temperature 0 is not deterministic; logprobs and the poor calibration of post-RLHF models.
The input/output price asymmetry, TTFT versus TPOT, why an agent loop is quadratic in turns, and effort as a dial.
Five examples, a binary check, a pass rate — deliberately early, so every later part is measurable.
System prompts, delimiters, and data/instruction separation — the first security lesson.
Selection, ordering, and format contagion: the model copies your examples' mistakes.
Wei et al. and Kojima et al. as history; self-consistency as the first test-time-compute pattern; why prescriptive CoT now often hurts.
JSON Schema, strict tools, and grammar/FSM token masking; syntax guarantees are not semantic ones.
Write, select, compress, isolate; compaction, JIT retrieval, structured notes; context rot and lost-in-the-middle; long context versus retrieval.
Prefix matching, stable-content-first ordering, breakpoints, silent invalidators, and measuring cache_read_input_tokens.
Freshness, privacy, citations, ACLs and cost — not 'the model doesn't know things'; BM25 as the brutally strong baseline.
What a vector means, cosine versus relevance, query/document asymmetry, Matryoshka truncation, and MTEB's overfitting problem.
Fixed, recursive, semantic, parent-doc and RAPTOR chunking; boundary effects; and contextual retrieval's measured failure reductions.
Flat, IVF, PQ, HNSW, ScaNN and DiskANN; recall, QPS and memory as a four-way tradeoff; and why filtered ANN is genuinely hard.
RRF versus score normalisation, SPLADE, cross-encoders and late interaction; retrieve wide, rerank narrow.
HyDE, multi-query and step-back; Self-RAG and CRAG; GraphRAG's global questions; ColPali; grep versus embeddings for code; and the long-context argument as an argument.
Every term in the volume linked back to the part that introduces it, plus a retrieval-architecture selection table and a RAG failure-mode to metric table.
Agents in Action, Interactively
The five workflow patterns, the complexity/value/viability/cost-of-error gate, and why most 'agents' should be workflows.
The five workflow patterns, the complexity/value/viability/cost-of-error gate, and why most 'agents' should be workflows.
Schemas, tool_use/tool_result, parallel calls, and errors as data; few non-overlapping tools, token-efficient returns, and error messages written for the model.
ReAct, Plan-and-Execute, Reflexion, ToT and LATS — what survived to production, and why the ideas outlived the algorithms.
The protocol and its stateless core; tool-definition bloat; SKILL.md's three loading stages; and why the context window is not a data bus.
Working, episodic, semantic and procedural memory; compaction versus context editing versus external notes; MemGPT/Letta, Mem0 and temporal graphs; compaction loss is adversarial.
Orchestrator-worker against single-thread; Anthropic's research system against Cognition's objection, and the read-heavy-and-independent versus shared-mutable-state reconciliation.
Context folding, step budgets, loop detection, sandboxing, permission gates, computer use, deep research as the composite case, and the harness as a source of variance.
Look at your data; open and axial coding; the failure taxonomy and its Pareto chart; build your own trace viewer.
Golden sets, coverage matrices, held-out test sets, code checks before judges, regression gating and significance.
Pointwise versus pairwise, rubrics from the taxonomy, position/verbosity/self-preference bias and its mitigation, validating the judge against humans, and IAA first.
Component-wise metrics, faithfulness and context precision/recall; end-state versus trajectory, pass^k, cost per completed task; SWE-bench Verified, τ²-bench, GAIA and OSWorld — and why public numbers don't transfer.
OTel GenAI semantics, logging retrieved context and prompt versions, A/B power, implicit signals, and cost and latency on the same dashboard.
The decision tree from prompt to RAG to eval to tune to RL; knowledge versus behaviour; the upgrade treadmill; your asset is the dataset, not the weights.
Curation, synthetic data and its diversity collapse, rejection sampling, LoRA/QLoRA operations, DPO/KTO/ORPO, RLVR and GRPO, reward hacking, and adapter serving.
Direct versus indirect injection; private data, untrusted content and an exfiltration channel; why no system prompt fixes it.
The six design patterns, capability-based information flow, dual-LLM; rails and classifiers as defence in depth; PII redaction before logging; GCG, PAIR and AgentDojo in CI.
Timeouts, backoff, circuit breakers and degraded modes; the ordered cost-lever list; three cache layers and semantic caching's correctness cost; routing and when not to cascade; streaming and speculative UX.
Prompts as versioned artifacts, canaries and rollback, risk-tiered autonomy, UX for uncertainty, and capturing user edits as signal.
Multimodal and voice and the ~800 ms budget; the small-model/on-device shift; A2A/AP2/x402 and what protocols cannot express; reliability compounding; how to keep up without chasing.
Every term in the volume linked back to the part that introduces it, plus the prompt to RAG to tune decision table, an eval-metric-per-failure-mode table, an injection-defense-pattern table, and a dated benchmark reference card.
Generative Media, Interactively
Sampling from a distribution you never see: a 2D dataset you draw yourself, and the line between memorising it and modelling it.
Sampling from a distribution you never see: a 2D dataset you draw yourself, and the line between memorising it and modelling it.
The forward process on a procedurally-generated image, q(x_t|x_0) in closed form, and what ᾱ, SNR and the schedule choice actually do.
ε-, x₀- and v-prediction, and a real tiny denoiser trained in the browser with its loss curve live.
The reverse loop from noise to sample; ancestral DDPM against deterministic DDIM, and a step-count slider that shows the trade.
The score field as arrows over a density, Langevin dynamics, the SMLD to DDPM equivalence and the VE versus VP SDEs.
Curved paths against straight ones, the velocity-field view of an ODE, and why the frontier trainers made flow matching the default.
The VAE compressor that shrinks an image before it is diffused, the cost of an eight-times downsample, and the pixel versus latent budget.
The resolution pyramid and its skips, residual blocks, where self- and cross-attention sit, and interactive parameter counts.
Patching a latent into tokens, plain transformer blocks with adaLN-Zero conditioning, and U-Net against DiT.
Text encoders and cross- against joint attention, CFG as extrapolation, and what guidance scale costs in diversity.
Euler, Heun and multistep solvers against the ancestral sampler, timestep spacing as a design choice, and step budget against error.
Progressive, consistency, adversarial and shortcut distillation, and what collapsing fifty steps into one or four costs in diversity.
LoRA and DreamBooth, ControlNet and IP-Adapter, inpainting and SDEdit editing, and where along the trajectory each intervention bites.
The causal 3-D VAE, spatiotemporal patches, 3-D against factorised attention, and the token budget a clip really costs.
Image and keyframe conditioning, temporal cascades, autoregressive rollout against full-sequence diffusion, drift, and the world-model claim.
A waveform, its sample rate, the STFT and the mel scale, and the token-rate arithmetic that makes raw waveform generation hopeless.
The encoder, residual vector quantiser and decoder that turn sound into a few hundred discrete tokens a second.
From alignment and duration models to codec language models and two-stage synthesis, with cloning and streaming first-chunk latency.
CTC, RNN-T and Whisper, why streaming is hard and the two-pass fix, then voice activity, barge-in and the realtime duplex budget.
FID, FVD and CLIPScore and what each misses, preference models and Elo arenas, memorisation, and provenance under the EU AI Act.
Every term linked back to the part that introduces it, plus the model, sampler and formula tables.
Multimodal Models, Interactively
Patching an image into a sequence, position embeddings and the class token, and what a vision transformer gives up against a convolution.
Patching an image into a sequence, position embeddings and the class token, and what a vision transformer gives up against a convolution.
InfoNCE on the similarity matrix, the temperature and batch size it needs, and SigLIP's pairwise sigmoid alternative.
Zero-shot classification, retrieval and prompt templates, and the compositionality, counting and typographic failures the space inherits.
Audio, video and 3D encoders aligned into one space, what a hub modality binds, and what aligned does and does not mean.
Linear and MLP projectors, the Q-Former and cross-attention bridges, where image tokens sit, and what each option trains.
Fixed patches, AnyRes tiling and native dynamic resolution, the merge that compresses tokens, and the image's prefill cost.
The staged recipe, freezing schedules and data mixtures, what breaks along the way, and preference tuning on multimodal pairs.
Objects that were never there, attention sinks on image tokens, the language prior overriding pixels, and the benchmarks that measure it.
Dropping the bridge between towers, native multimodal pretraining, modality embedders into one stream, and modality-aware experts.
Frame sampling against a token budget, temporal position encoding, streaming memory, and the KV arithmetic a clip implies.
Speech as an encoder's output or as native audio tokens in the trunk, speech instruction following, and paralinguistics.
Interleaved token streams, the thinker-and-talker split, listening while speaking, barge-in, and the realtime budget.
One model that reads and makes images, the tokenizer choice that decides whether it works, and three generation orders.
Detection and segmentation as coordinate tokens, promptable segmentation, referring expressions, and grounding an interface.
Actions as text tokens, a flow-matching action expert, chunking, and the latent world models that predict what happens next.
Encoder disaggregation and caching, the prefill-and-decode inversion an image causes, safety, and the open problems.
Every term linked back to the part that introduces it, plus the architecture-era, token-budget and evaluation tables.