Serving mixture-of-experts at scale
A mixture-of-experts model is a bet that you can buy quality with capacity rather than compute. DeepSeek-V3 holds 671 billion parameters but reads only 37 billion per token; Kimi K2 holds a trillion and reads 32 billion. On paper that is a dense-quality model at a fraction of the FLOPs. In practice it breaks the two assumptions every dense serving stack is built on: that the weights you hold are the weights you read, and that a batch of tokens does the same work as one token. This part is about the machinery that makes the bet pay — routing, all-to-all, hot experts, EPLB — and the rack-scale interconnect it demands.
Active vs total parameters
Two numbers, not one
A dense model has one parameter count because every parameter participates in every token. An MoE has two: the total parameters you must store, and the active parameters the router actually reads for a given token. DeepSeek-V3 activates 8 of 256 routed experts plus one shared expert, so 37B of its 671B parameters move per token — 5.5%. Pick a model below and compare its compute against two hypothetical dense baselines: one with the same total parameters (same capacity, 18× the FLOPs) and one with the same active parameters (same FLOPs, a fraction of the capacity). The training guide's transformer chapter introduced the router and top-k; this page is about the cost of running them.
Bars are FLOPs per token. The two dense baselines bracket the MoE's trade.
The memory paradox
Capacity-bound, not bandwidth-bound
A dense 70B model in BF16 needs ~140 GB of weights and reads ~140 GB per token. DeepSeek-V3 needs ~1,342 GB to hold but reads only ~74 GB per token. That inverts the usual deficit: MoE serving is limited first by the capacity to keep every expert resident — across many GPUs — and only second by the bandwidth to stream the active subset. It also removes the clean roofline: a decode step's arithmetic intensity depends on the batch (next section), and the cost of an expert being cold (not resident, fetched from host or disk) is a stall, not a slow read. The practical consequence is that MoE deployments are laid out for capacity first: they pile experts into as many GPUs as the model needs, then optimise communication.
Routing and top-k
A token picks a few experts
The router is a learned linear layer that scores all \(E\) experts for a token; the top \(k\) are selected, their outputs are combined by a softmax-weighted sum, and the rest are not computed. DeepSeek-V3 uses \(E=256\), \(k=8\), plus one always-on shared expert; Kimi K2 uses 384 routed experts, 8 active plus 1 shared. Routing is token-level and per-layer, so a sequence's tokens fan out across different experts every layer. Two serving consequences follow immediately. First, the per-token FLOPs are roughly \(2 \times \text{active}\) while the weights held are \(2 \times \text{total}\), which is why quantization of experts buys capacity more than it buys speed. Second, routing quality and routing balance are separate problems: a router can be accurate and still send a disproportionate share of tokens to a handful of experts, which is the subject of the next two sections.
All-to-all dispatch and combine
The collective that replaces all-reduce
With expert parallelism, each device holds a subset of experts. For every MoE layer a token's hidden state is dispatched to the \(k\) devices holding its chosen experts, the expert MLPs run locally, and the results are combined back. That is an all-to-all — every device sends to every other — not the ring all-reduce of tensor parallelism. Its volume per layer is \(k\) sends and \(k\) receives of the hidden state per token, plus index and scale metadata. Across 61 layers at batch 256 with hidden 7,168, the round trip is several gigabytes per step. On NVLink that is a few milliseconds; on Ethernet it is the majority of the step, and the wider the EP, the more nodes the traffic crosses. Toggle the fabric to see the crossover that motivates Epoxy/DeepEP's use of NVLink and RDMA directly.
Stacked step time: compute plus the dispatch/combine all-to-all. The share is what wide EP forces you to care about.
Hot experts
Load is never uniform
Routing is learned from data, and data has structure: some experts specialise in punctuation, code, numbers or a language, and they are chosen far more often than the rest. DeepSeek notes "inherently high-load experts," and the widely repeated rule of thumb is that the top 3–5% of experts absorb 30–50% of tokens. The heatmap below routes 2,000 seeded tokens through a 256-expert top-8 router and plots load per expert over time. Two facts matter operationally. The distribution is skewed (the Gini coefficient and the top-10 share are computed live), and the step time is set by the busiest expert, not the average — a straggler stalls every other device waiting for the combine. Perfectly balanced routing is not a goal you reach by training alone; it is an inference-time scheduling problem.
Rows are experts sorted by total load; columns are 50-token time bins. Play to reveal the timeline.
EPLB: expert-parallel load balancing
Replicate the hot ones
DeepSeek's open-sourced expert-parallel load balancer takes a different route from changing the router: it observes per-expert load and replicates the hot experts onto devices that have spare capacity, so the same expert runs on two or more GPUs and the router can split its traffic. The cost is memory — each replica is a full expert — and the benefit is a lower maximum per-device load and a shorter step. The commonly cited recovery is around 20%. Toggle replication below: watch the per-device load bars flatten and the extra expert memory appear.
Per-device load. The height of the tallest bar is the step time.
Wide EP and DeepEP
Scaling the collective, not the model
DeepSeek's production decode runs EP144 — experts spread across 144 GPUs, which is 18 nodes of 8. At that width the all-to-all crosses node boundaries constantly, so the communication library matters more than the model. DeepEP is a communication library built for exactly this: high-throughput kernels for dispatch and combine, low-latency kernels for decode, group-limited routing to cap how far a token can travel, and direct RDMA over NVLink and InfiniBand to avoid a copy through host memory. This is the same pattern Part 12 described for tensor parallelism, pushed to a different collective: keep the communication on the fastest fabric you have, and make the fabric one big domain. The economics only work because wide EP serves a huge batch with a fixed set of weights — the disaggregated pools of Part 13 are the natural home for it.
Batch size changes everything
Sparse at batch 1, dense at batch 512
At batch 1 a token touches 8 of 256 experts — about 6% of the weights, once attention, the shared expert and the router are counted. At batch 512 the union of chosen experts is almost every expert, so the step reads ~100% of them even though each token still computes only 8. Sparsity is a per-token property; at serving batch sizes the system is effectively doing dense reads of all the expert weights, just with a smaller per-token compute bill. That has three consequences: arithmetic intensity rises with batch, so MoE decode is less bandwidth-starved than dense decode; capacity still has to cover all experts; and the crossover batch at which the union saturates is the batch you want to run at, because below it you are paying the communication of EP without amortising the weight reads. Drag the slider to see the union grow.
Fraction of total expert weights read per step as the batch of routed tokens grows.
Why rack-scale exists
The interconnect is the design
Put the previous sections together and the rack is not a packaging choice — it is the algorithm. Wide EP needs a collective that is cheap at 100+ participants; hot-expert replication needs spare capacity next door; batch saturation needs a big enough domain to route hundreds of tokens without crossing a slow link. GB200 NVL72 puts 72 GPUs in a single NVLink domain at ~1.8 TB/s per GPU and ~130 TB/s aggregate, which is roughly the scale at which all-to-all stops being the bottleneck. That is why MoE serving and rack-scale NVLink appeared together, and why a MoE deployment's first infrastructure question is not "how many GPUs" but "how big is one domain."
Cheat sheet
| Concept | What to remember |
|---|---|
| Two parameter counts | Total = capacity you hold; active = weights read per token. DeepSeek-V3: 671B / 37B |
| Memory paradox | MoE is capacity-bound first: holds 1,342 GB, reads ~74 GB/token |
| Top-k routing | 256 experts, top-8 + 1 shared for DeepSeek-V3; Kimi K2 uses 384 routed, 8 + 1 |
| All-to-all | Dispatch + combine replaces all-reduce; volume ∝ k × hidden × layers × batch |
| Hot experts | Top 3–5% absorb 30–50% of tokens; step time is the max load, not the mean |
| EPLB | Replicate hot experts; ~20% recovery at the cost of expert-sized memory |
| Wide EP | EP144 decode for DeepSeek-V3; DeepEP for NVLink/RDMA all-to-all |
| Batch effect | ~6% of weights at B=1, ~100% at B=512 — sparsity is per-token, not per-step |
Further reading
- DeepSeek-AI, "DeepSeek-V3 Technical Report" (2024) — 256 routed + 1 shared expert, top-8, auxiliary-loss-free balancing.
- DeepSeek-AI, "DeepSeek-V3/R1 Inference System Overview" (2025-02-28) — EP32 prefill, EP144 decode.
- DeepSeek-AI, EPLB: Expert Parallel Load Balancer (2025).
- DeepSeek-AI, DeepEP — high-throughput and low-latency expert-parallel communication.
- Lepikhin et al., "GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding" (2020) — the top-k MoE serving pattern.