Parallelism for inference: TP, PP, EP, DP-attention, CP
A 70B model does not fit on one 80 GB GPU, and a 671B model with a million-token KV cache does not fit on eight. Serving therefore shards a single model across devices — but the sharding that trains a model and the sharding that serves it are different problems. Training optimises time-to-convergence across a large global batch and can hide communication behind gradient accumulation; inference must hit a per-request latency SLO while every collective sits on the critical path. This part builds the five axes — tensor, pipeline, expert, data and context parallelism — and the constraint that organises all of them: keep the expensive collectives inside the NVLink domain.
Inference parallelism is not training parallelism
Same model, opposite objective
The training guide's inference chapter established the two facts this whole series runs on: a decode step reads every active weight from HBM to produce one token, and prefill vs decode are two different bottlenecks. What that chapter did not need is the parallelism — one GPU, one model. Here we generalise.
Training parallelism exists to make a global batch of thousands of sequences fit and converge faster. Data parallelism (DP) with ZeRO shards optimizer states, gradients and parameters; tensor parallelism (TP) splits a layer's matrices so a single layer's compute is spread across devices; pipeline parallelism (PP) splits layers into stages. Communication is usually overlapped with compute or amortised over accumulation steps, and a few milliseconds of all-reduce is invisible against a days-long run.
Inference has no accumulation to hide behind and a latency SLO to meet per token. The decode step is memory-bandwidth-bound with an arithmetic intensity near 1–2 FLOPs/byte, so the way to make it faster is to read less per step — or to read the same weights while serving more sequences. That single fact explains the ranking of the axes: DP (replicate the model, split the requests) is almost free and raises throughput without touching single-request latency; TP is expensive but the only way to shrink the weight read on the critical path; PP is cheap on bandwidth but adds bubbles and stage hops; EP is a bandwidth play for sparse models; CP is a reach tool for sequences too long for one device.
Tensor parallelism: two all-reduces per layer
The only way to cut the weight read
TP splits each weight matrix. In the Megatron scheme the attention and MLP blocks each do a column-parallel projection, a local nonlinearity, and a row-parallel projection, which needs an all-reduce of the partial sums. That makes two all-reduces per transformer layer, every step, on the critical path. A ring all-reduce moves roughly
where \(S\) is the activation size, \(N\) the TP degree, \(B_{\text{link}}\) the link bandwidth and \(\ell\) the per-hop latency. Inside NVLink this is microseconds; across nodes it is tens of milliseconds. The demo below sweeps TP from 2 to 32 on a 70B-class model at batch 256 and plots scaling efficiency — speed-up divided by TP. Watch it flatten by TP=8, then reverse: past the NVLink domain the all-reduce cost grows faster than the weight shard shrinks, so adding GPUs makes the step slower.
Efficiency is speed-up ÷ TP, measured against TP=2. The dashed rule marks the 8-GPU NVLink domain.
Pipeline parallelism and the bubble
Cheap on bandwidth, expensive on latency
PP splits the transformer's layers into \(P\) stages and streams microbatches through them. Only activations cross a stage boundary — a point-to-point send of \(batch \times hidden\) elements, far smaller than an all-reduce — so PP is attractive across a slow link. Its cost is the pipeline bubble: while the pipe fills and drains, stages sit idle. For \(M\) microbatches and \(P\) stages the unavoidable idle fraction is
Training hides the bubble with gradient accumulation. Inference's equivalent is continuous batching: the queue supplies a stream of sequences, so \(M\) is effectively the number of concurrent requests. With \(M=128\) and \(P=4\), the bubble is 3/131 ≈ 2.3%; with \(M=4\) it is 43%. The toggle below swaps between a thin batch and a continuously-batched pool. This is why disaggregated or high-QPS inference can afford PP across nodes, while a single interactive stream usually cannot.
Each row is a pipeline stage; each block is one microbatch's forward pass. Grey is the bubble.
Expert parallelism (preview of Part 14)
Shard the experts, not the matrices
A mixture-of-experts layer is a router plus \(E\) independent feed-forward experts. EP places different experts on different devices, so a token's hidden state must be sent to whichever devices host its chosen experts, then sent back — an all-to-all rather than an all-reduce. DeepSeek-V3 uses 256 routed experts plus one shared expert and activates 8, and its production engine runs expert parallelism as wide as EP32 for prefill and EP144 for decode. That width is only viable because the all-to-all stays inside a rack-scale NVLink domain; over Ethernet it would dominate the step. Part 14 builds the full cost model and the EPLB rebalancer. For now: EP buys capacity without spending bandwidth on dense weights, at the price of a new collective.
DP and DP-attention for MLA
Replicas are free until the KV cache is not
Data parallelism keeps a full model replica per group and sends each replica a different set of requests. In inference there are no gradients to synchronise, so pure DP costs no collective at all — it is the cheapest way to raise system throughput, and the first axis to spend before touching TP. A DP replica serving \(B\) sequences produces \(B\) tokens per step, so throughput scales with the replica count while single-request latency is unchanged.
The catch is memory. GQA and MLA shrink the per-token KV cache, but at long context it can still exceed the weights. That motivates DP-attention: instead of splitting the attention weight matrices with TP (and paying an all-reduce around attention), replicate the attention computation across DP ranks and shard the KV cache along the sequence or head dimension. For MLA, whose cache is a compressed latent plus a small RoPE key, the allocation across ranks is nearly free of compute cost, and decode's attention reads become a partitioned gather. The result: bigger effective batch and less KV per device, without a second TP collective per layer.
Context parallelism (preview of Part 15)
When one sequence no longer fits
TP, PP and DP all leave the sequence axis intact. Context parallelism (CP) splits the tokens of a single sequence across devices, so a 200K-token prompt's attention is computed in pieces and combined with ring-style point-to-point communication. CP is exactly the training technique called sequence parallelism, transplanted to a workload where the "batch" may be one very long request. It is the last resort: it adds communication proportional to sequence length and is only worth it when neither KV sharding nor memory tiers can hold the context — the territory of Part 15.
The search space: every way to spend 32 GPUs
Interactive topology explorer
A real deployment is a factorisation \(G = TP \times PP \times DP\) (plus EP and CP when present). Not every factorisation is legal or fast: the weights and KV cache must fit, TP must stay inside the NVLink domain, and a short batch makes the pipeline bubble large. The explorer below lays out 32 GPUs — blue edges are TP all-reduce links, magenta edges are PP activation hops, rows are DP replicas. Change TP, PP, DP or the per-replica batch and read the step time, system throughput, memory per GPU and the verdict.
Keep TP inside NVLink; replicas before shards
Reading the search space
The Pareto sweep below evaluates every factorisation of 32 GPUs at the current batch. The horizontal axis is step time; the vertical axis is system throughput. Points up and to the left dominate. Click any legal point to load it into the explorer above.
Circled points form the Pareto frontier. Red points fail a constraint; click a point to load it.
The shape of that cloud is the whole lesson. The best throughput usually comes from small TP (2–8), modest PP, and as much DP as fits — because DP is free and TP's all-reduce is not. When a single request must be fast, TP grows until the weight read fits the latency budget, and it stops at the node boundary. PP appears only when a model is too large for one node even with TP=8, and then only when continuous batching supplies enough microbatches to keep the bubble under a few percent. EP replaces TP for sparse layers and is the reason rack-scale domains exist.
| Axis | Splits | Collective | Use it when |
|---|---|---|---|
| DP | Requests / replicas | None in inference | Always first — raises throughput for free |
| TP | Weight matrices | 2 all-reduces / layer | Weights or latency demand it, and TP ≤ NVLink domain |
| PP | Layer blocks | Point-to-point activations | Model spans nodes; M is large so the bubble is small |
| EP | Experts | All-to-all dispatch/combine | Sparse MoE; rack-scale NVLink makes it affordable |
| DP-attention | KV cache / attention | Gather, no all-reduce | MLA/GQA cache dominates memory |
| CP | Sequence tokens | Ring point-to-point | One sequence exceeds a device's KV budget |
Cheat sheet
| Concept | What to remember |
|---|---|
| Two all-reduces per layer | TP's attention and MLP blocks each need one all-reduce of partial sums, every step |
| Ring all-reduce volume | \(2(N-1)/N \times S/B_{\text{link}}\), plus \((N-1)\ell\) latency |
| NVLink cliff | 900 GB/s intra-node vs ~50 GB/s inter-node on H100-class — 18× |
| Pipeline bubble | \((P-1)/(M+P-1)\); continuous batching raises M and shrinks it |
| DP is free in inference | No gradients, so replicas need no collective; spend DP before TP |
| DP-attention | Replicate attention, shard the KV cache — avoids a TP all-reduce on MLA |
| Wide EP | EP32 prefill / EP144 decode for DeepSeek-V3, only on a rack-scale domain |
Further reading
- Shoeybi et al., "Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism" (2019) — the two-all-reduce TP scheme.
- Narayanan et al., "Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM" (2021) — interleaved PP and bubble arithmetic.
- Liu et al., "Ring Attention with Blockwise Transformers for Near-Infinite Context" (2023) — the CP primitive.
- DeepSeek-AI, "DeepSeek-V3 Technical Report" (2024) — 256 routed experts, EP32/EP144 deployment.
- NVIDIA, GB200 NVL72 — 72 GPUs, ~1.8 TB/s NVLink per GPU, one 130 TB/s domain.