Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

Active vs total parameters

Two numbers, not one

A dense model has one parameter count because every parameter participates in every token. An MoE has two: the total parameters you must store, and the active parameters the router actually reads for a given token. DeepSeek-V3 activates 8 of 256 routed experts plus one shared expert, so 37B of its 671B parameters move per token — 5.5%. Pick a model below and compare its compute against two hypothetical dense baselines: one with the same total parameters (same capacity, 18× the FLOPs) and one with the same active parameters (same FLOPs, a fraction of the capacity). The training guide's transformer chapter introduced the router and top-k; this page is about the cost of running them.

Bars are FLOPs per token. The two dense baselines bracket the MoE's trade.

2

The memory paradox

Capacity-bound, not bandwidth-bound

A dense 70B model in BF16 needs ~140 GB of weights and reads ~140 GB per token. DeepSeek-V3 needs ~1,342 GB to hold but reads only ~74 GB per token. That inverts the usual deficit: MoE serving is limited first by the capacity to keep every expert resident — across many GPUs — and only second by the bandwidth to stream the active subset. It also removes the clean roofline: a decode step's arithmetic intensity depends on the batch (next section), and the cost of an expert being cold (not resident, fetched from host or disk) is a stall, not a slow read. The practical consequence is that MoE deployments are laid out for capacity first: they pile experts into as many GPUs as the model needs, then optimise communication.

3

Routing and top-k

A token picks a few experts

The router is a learned linear layer that scores all \(E\) experts for a token; the top \(k\) are selected, their outputs are combined by a softmax-weighted sum, and the rest are not computed. DeepSeek-V3 uses \(E=256\), \(k=8\), plus one always-on shared expert; Kimi K2 uses 384 routed experts, 8 active plus 1 shared. Routing is token-level and per-layer, so a sequence's tokens fan out across different experts every layer. Two serving consequences follow immediately. First, the per-token FLOPs are roughly \(2 \times \text{active}\) while the weights held are \(2 \times \text{total}\), which is why quantization of experts buys capacity more than it buys speed. Second, routing quality and routing balance are separate problems: a router can be accurate and still send a disproportionate share of tokens to a handful of experts, which is the subject of the next two sections.

4

All-to-all dispatch and combine

The collective that replaces all-reduce

With expert parallelism, each device holds a subset of experts. For every MoE layer a token's hidden state is dispatched to the \(k\) devices holding its chosen experts, the expert MLPs run locally, and the results are combined back. That is an all-to-all — every device sends to every other — not the ring all-reduce of tensor parallelism. Its volume per layer is \(k\) sends and \(k\) receives of the hidden state per token, plus index and scale metadata. Across 61 layers at batch 256 with hidden 7,168, the round trip is several gigabytes per step. On NVLink that is a few milliseconds; on Ethernet it is the majority of the step, and the wider the EP, the more nodes the traffic crosses. Toggle the fabric to see the crossover that motivates Epoxy/DeepEP's use of NVLink and RDMA directly.

Stacked step time: compute plus the dispatch/combine all-to-all. The share is what wide EP forces you to care about.

5

Hot experts

Load is never uniform

Routing is learned from data, and data has structure: some experts specialise in punctuation, code, numbers or a language, and they are chosen far more often than the rest. DeepSeek notes "inherently high-load experts," and the widely repeated rule of thumb is that the top 3–5% of experts absorb 30–50% of tokens. The heatmap below routes 2,000 seeded tokens through a 256-expert top-8 router and plots load per expert over time. Two facts matter operationally. The distribution is skewed (the Gini coefficient and the top-10 share are computed live), and the step time is set by the busiest expert, not the average — a straggler stalls every other device waiting for the combine. Perfectly balanced routing is not a goal you reach by training alone; it is an inference-time scheduling problem.

Rows are experts sorted by total load; columns are 50-token time bins. Play to reveal the timeline.

6

EPLB: expert-parallel load balancing

Replicate the hot ones

DeepSeek's open-sourced expert-parallel load balancer takes a different route from changing the router: it observes per-expert load and replicates the hot experts onto devices that have spare capacity, so the same expert runs on two or more GPUs and the router can split its traffic. The cost is memory — each replica is a full expert — and the benefit is a lower maximum per-device load and a shorter step. The commonly cited recovery is around 20%. Toggle replication below: watch the per-device load bars flatten and the extra expert memory appear.

Per-device load. The height of the tallest bar is the step time.

7

Wide EP and DeepEP

Scaling the collective, not the model

DeepSeek's production decode runs EP144 — experts spread across 144 GPUs, which is 18 nodes of 8. At that width the all-to-all crosses node boundaries constantly, so the communication library matters more than the model. DeepEP is a communication library built for exactly this: high-throughput kernels for dispatch and combine, low-latency kernels for decode, group-limited routing to cap how far a token can travel, and direct RDMA over NVLink and InfiniBand to avoid a copy through host memory. This is the same pattern Part 12 described for tensor parallelism, pushed to a different collective: keep the communication on the fastest fabric you have, and make the fabric one big domain. The economics only work because wide EP serves a huge batch with a fixed set of weights — the disaggregated pools of Part 13 are the natural home for it.

8

Batch size changes everything

Sparse at batch 1, dense at batch 512

At batch 1 a token touches 8 of 256 experts — about 6% of the weights, once attention, the shared expert and the router are counted. At batch 512 the union of chosen experts is almost every expert, so the step reads ~100% of them even though each token still computes only 8. Sparsity is a per-token property; at serving batch sizes the system is effectively doing dense reads of all the expert weights, just with a smaller per-token compute bill. That has three consequences: arithmetic intensity rises with batch, so MoE decode is less bandwidth-starved than dense decode; capacity still has to cover all experts; and the crossover batch at which the union saturates is the batch you want to run at, because below it you are paying the communication of EP without amortising the weight reads. Drag the slider to see the union grow.

Fraction of total expert weights read per step as the batch of routed tokens grows.

9

Why rack-scale exists

The interconnect is the design

Put the previous sections together and the rack is not a packaging choice — it is the algorithm. Wide EP needs a collective that is cheap at 100+ participants; hot-expert replication needs spare capacity next door; batch saturation needs a big enough domain to route hundreds of tokens without crossing a slow link. GB200 NVL72 puts 72 GPUs in a single NVLink domain at ~1.8 TB/s per GPU and ~130 TB/s aggregate, which is roughly the scale at which all-to-all stops being the bottleneck. That is why MoE serving and rack-scale NVLink appeared together, and why a MoE deployment's first infrastructure question is not "how many GPUs" but "how big is one domain."

📌 The MoE serving checklist: experts resident (capacity) → router balanced (EPLB) → all-to-all on the fast fabric (wide EP + DeepEP) → batch large enough to amortise (continuous batching). Miss any one and the density advantage evaporates.

Cheat sheet

ConceptWhat to remember
Two parameter countsTotal = capacity you hold; active = weights read per token. DeepSeek-V3: 671B / 37B
Memory paradoxMoE is capacity-bound first: holds 1,342 GB, reads ~74 GB/token
Top-k routing256 experts, top-8 + 1 shared for DeepSeek-V3; Kimi K2 uses 384 routed, 8 + 1
All-to-allDispatch + combine replaces all-reduce; volume ∝ k × hidden × layers × batch
Hot expertsTop 3–5% absorb 30–50% of tokens; step time is the max load, not the mean
EPLBReplicate hot experts; ~20% recovery at the cost of expert-sized memory
Wide EPEP144 decode for DeepSeek-V3; DeepEP for NVLink/RDMA all-to-all
Batch effect~6% of weights at B=1, ~100% at B=512 — sparsity is per-token, not per-step

Further reading

?

Check your understanding

0/5 answered