Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

Inference parallelism is not training parallelism

Same model, opposite objective

The training guide's inference chapter established the two facts this whole series runs on: a decode step reads every active weight from HBM to produce one token, and prefill vs decode are two different bottlenecks. What that chapter did not need is the parallelism — one GPU, one model. Here we generalise.

Training parallelism exists to make a global batch of thousands of sequences fit and converge faster. Data parallelism (DP) with ZeRO shards optimizer states, gradients and parameters; tensor parallelism (TP) splits a layer's matrices so a single layer's compute is spread across devices; pipeline parallelism (PP) splits layers into stages. Communication is usually overlapped with compute or amortised over accumulation steps, and a few milliseconds of all-reduce is invisible against a days-long run.

Inference has no accumulation to hide behind and a latency SLO to meet per token. The decode step is memory-bandwidth-bound with an arithmetic intensity near 1–2 FLOPs/byte, so the way to make it faster is to read less per step — or to read the same weights while serving more sequences. That single fact explains the ranking of the axes: DP (replicate the model, split the requests) is almost free and raises throughput without touching single-request latency; TP is expensive but the only way to shrink the weight read on the critical path; PP is cheap on bandwidth but adds bubbles and stage hops; EP is a bandwidth play for sparse models; CP is a reach tool for sequences too long for one device.

💡 The organising constraint: every collective you put on the critical path must be cheap. On an H100, NVLink moves 900 GB/s and an InfiniBand-class inter-node link moves roughly 50 GB/s — an 18× cliff. Almost every real deployment layout is an attempt to keep TP (and later, hot all-to-all) on the fast side of that cliff.
2

Tensor parallelism: two all-reduces per layer

The only way to cut the weight read

TP splits each weight matrix. In the Megatron scheme the attention and MLP blocks each do a column-parallel projection, a local nonlinearity, and a row-parallel projection, which needs an all-reduce of the partial sums. That makes two all-reduces per transformer layer, every step, on the critical path. A ring all-reduce moves roughly

$$ T_{\text{AR}} \approx \underbrace{2\frac{N-1}{N}}_{\text{volume}}\frac{S}{B_{\text{link}}} + \underbrace{(N-1)\ell}_{\text{latency}}, $$

where \(S\) is the activation size, \(N\) the TP degree, \(B_{\text{link}}\) the link bandwidth and \(\ell\) the per-hop latency. Inside NVLink this is microseconds; across nodes it is tens of milliseconds. The demo below sweeps TP from 2 to 32 on a 70B-class model at batch 256 and plots scaling efficiency — speed-up divided by TP. Watch it flatten by TP=8, then reverse: past the NVLink domain the all-reduce cost grows faster than the weight shard shrinks, so adding GPUs makes the step slower.

Efficiency is speed-up ÷ TP, measured against TP=2. The dashed rule marks the 8-GPU NVLink domain.

3

Pipeline parallelism and the bubble

Cheap on bandwidth, expensive on latency

PP splits the transformer's layers into \(P\) stages and streams microbatches through them. Only activations cross a stage boundary — a point-to-point send of \(batch \times hidden\) elements, far smaller than an all-reduce — so PP is attractive across a slow link. Its cost is the pipeline bubble: while the pipe fills and drains, stages sit idle. For \(M\) microbatches and \(P\) stages the unavoidable idle fraction is

$$ \text{bubble} = \frac{P-1}{M+P-1}. $$

Training hides the bubble with gradient accumulation. Inference's equivalent is continuous batching: the queue supplies a stream of sequences, so \(M\) is effectively the number of concurrent requests. With \(M=128\) and \(P=4\), the bubble is 3/131 ≈ 2.3%; with \(M=4\) it is 43%. The toggle below swaps between a thin batch and a continuously-batched pool. This is why disaggregated or high-QPS inference can afford PP across nodes, while a single interactive stream usually cannot.

Each row is a pipeline stage; each block is one microbatch's forward pass. Grey is the bubble.

4

Expert parallelism (preview of Part 14)

Shard the experts, not the matrices

A mixture-of-experts layer is a router plus \(E\) independent feed-forward experts. EP places different experts on different devices, so a token's hidden state must be sent to whichever devices host its chosen experts, then sent back — an all-to-all rather than an all-reduce. DeepSeek-V3 uses 256 routed experts plus one shared expert and activates 8, and its production engine runs expert parallelism as wide as EP32 for prefill and EP144 for decode. That width is only viable because the all-to-all stays inside a rack-scale NVLink domain; over Ethernet it would dominate the step. Part 14 builds the full cost model and the EPLB rebalancer. For now: EP buys capacity without spending bandwidth on dense weights, at the price of a new collective.

📌 Rule of thumb: dense models parallelise with TP inside a node; sparse models parallelise with EP across a rack. The two collectives — all-reduce and all-to-all — have different scaling laws.
5

DP and DP-attention for MLA

Replicas are free until the KV cache is not

Data parallelism keeps a full model replica per group and sends each replica a different set of requests. In inference there are no gradients to synchronise, so pure DP costs no collective at all — it is the cheapest way to raise system throughput, and the first axis to spend before touching TP. A DP replica serving \(B\) sequences produces \(B\) tokens per step, so throughput scales with the replica count while single-request latency is unchanged.

The catch is memory. GQA and MLA shrink the per-token KV cache, but at long context it can still exceed the weights. That motivates DP-attention: instead of splitting the attention weight matrices with TP (and paying an all-reduce around attention), replicate the attention computation across DP ranks and shard the KV cache along the sequence or head dimension. For MLA, whose cache is a compressed latent plus a small RoPE key, the allocation across ranks is nearly free of compute cost, and decode's attention reads become a partitioned gather. The result: bigger effective batch and less KV per device, without a second TP collective per layer.

6

Context parallelism (preview of Part 15)

When one sequence no longer fits

TP, PP and DP all leave the sequence axis intact. Context parallelism (CP) splits the tokens of a single sequence across devices, so a 200K-token prompt's attention is computed in pieces and combined with ring-style point-to-point communication. CP is exactly the training technique called sequence parallelism, transplanted to a workload where the "batch" may be one very long request. It is the last resort: it adds communication proportional to sequence length and is only worth it when neither KV sharding nor memory tiers can hold the context — the territory of Part 15.

⚠️ CP on the critical path is painful: ring attention passes KV blocks around the ring every layer, so its cost grows with sequence length and ring width. A long-context request served with CP is faster on TTFT only if the alternative is swapping KV to host memory or recomputing.

Cheat sheet

ConceptWhat to remember
Two all-reduces per layerTP's attention and MLP blocks each need one all-reduce of partial sums, every step
Ring all-reduce volume\(2(N-1)/N \times S/B_{\text{link}}\), plus \((N-1)\ell\) latency
NVLink cliff900 GB/s intra-node vs ~50 GB/s inter-node on H100-class — 18×
Pipeline bubble\((P-1)/(M+P-1)\); continuous batching raises M and shrinks it
DP is free in inferenceNo gradients, so replicas need no collective; spend DP before TP
DP-attentionReplicate attention, shard the KV cache — avoids a TP all-reduce on MLA
Wide EPEP32 prefill / EP144 decode for DeepSeek-V3, only on a rack-scale domain

Further reading

?

Check your understanding

0/5 answered