Prefill/decode disaggregation
Colocating prefill and decode on the same GPU is convenient and, at high load, structurally wrong. Prefill is a long compute-bound burst; decode is a steady memory-bound drip that millions of users feel as inter-token latency. Put them on one device and every prefill pauses every decode step. Disaggregation separates the two jobs into pools with different hardware, different batching policies, even different GPU counts — and pays for it by shipping the KV cache across the fabric. This part is about when that trade wins, how to size the pools, and what DeepSeek disclosed about running it at 608 billion input tokens a day.
One pool, two jobs
The ITL spike you can see
The training guide's inference chapter framed prefill as compute-bound and decode as bandwidth-bound. On a shared engine, those two workloads contend for the same SMs: while a 4,000-token prompt is prefilled, the running decode batch waits its turn, and a user watching tokens stream sees a stall. Part 3 named the metric — inter-token latency, or ITL — and this is the mechanism that makes its p99 ugly even when throughput looks healthy.
Chunked prefill (Part 6) mitigates this by slicing prompts into iteration-sized pieces, but the total work is unchanged: some iterations still carry prefill, and the decode step is longer. Disaggregation attacks the problem by routing each request's two phases to different GPUs. Prefill workers batch many prompts together, saturating compute with large matrix multiplies; decode workers run a steady, latency-optimised loop with a fixed batch. The prefill worker sends the resulting KV cache to the decode worker over the network, and the decode worker streams tokens back. The demo below replays both designs: the top timeline is colocated and shows a prefill spike every few iterations; the bottom keeps decode flat and exposes a small transfer segment.
Each bar is one engine iteration, stacked by job. Play to advance the timeline.
DistServe, Splitwise and Mooncake
Three papers, one architecture
DistServe (OSDI 2024) introduced the formulation and reported serving 7.4× more requests or 12.6× tighter SLO by splitting the phases; the widely quoted 20× is a rounded composite, not a single measurement. Splitwise (Microsoft) made the same split for a different reason — phase-specific hardware, since prefill wants compute and decode wants memory capacity. Mooncake (Moonshot/Kimi) went further and made the KV cache itself the centre of the system, with a KVCache-centric scheduler and a transfer engine, reporting up to 525% throughput in simulation and 75% more real requests. The same architecture now ships as first-class support in vLLM and SGLang via disaggregation connectors.
Sizing the pools
The input:output ratio decides everything
Pool sizing is an arithmetic question. Let \(R_p\) be prefill tokens/s per GPU and \(R_d\) decode tokens/s per GPU. For a workload with \(P\) input and \(O\) output tokens per request, the GPU-seconds needed per request are \(a = P/R_p\) for prefill and \(b = O/R_d\) for decode, so the throughput-balanced split is
Because decode is bandwidth-bound and batched, \(R_p/R_d\) is large — with representative values of 12,000 prefill and 800 decode tokens/s per GPU it is 15. That makes the optimal prefill share \(r/(r+15)\): a chat workload at \(r=4\) wants ~21% prefill, summarisation at \(r=20\) wants ~57%, and a reasoning workload at \(r=1/6\) wants ~1%. DeepSeek's disclosed ratio of 608B input to 168B output tokens (\(r=3.62\)) predicts ~19% prefill — and their deployment runs 32 prefill to 144 decode GPUs, or 18.2%. Use the presets to see each workload, then drag the split away from optimal and watch queues form.
Bars show utilisation of each pool relative to the throughput-balanced allocation; the marker is the optimal prefill share.
Moving KV: NIXL, Mooncake Transfer Engine, RDMA
The cost you pay for freedom
The transfer is the whole bargain. A prefill worker holds a KV cache of \(2\,L\,n_{kv}d_{head}T\) bytes per sequence; for a 70B GQA model that is about 0.33 MB per token, so a 4,000-token prompt is ~1.3 GB. Moving that over a 400 Gb/s (50 GB/s) RDMA link takes ~26 ms; over 100 GbE it takes ~105 ms; over NVLink it is under 3 ms. Transfer engines (NVIDIA's NIXL, Mooncake Transfer Engine, LMCache) exist to drive RDMA queues at line rate, register GPU memory once and pipeline chunk-by-chunk, so the copy overlaps the decode that follows. MLA models cut the payload ~4.7× — DeepSeek's latent is 70 KB/token versus ~328 KB for a same-geometry GQA cache — which is a large part of why disaggregation is practical for them.
The crossing point is the minimum prompt length at which relocating the KV beats recomputing prefill in place.
Layer-wise overlap
Don't wait for the whole cache
The naive transfer waits for prefill to finish, sends the entire cache, then starts decode. Real systems pipeline at layer granularity: the prefill worker streams layer \(i\)'s KV as soon as it is computed, and the decode worker begins prefilling its own layer \(i\) state and starts generating the first token as the tail of the cache arrives. The critical path becomes the transfer of the last layer, not the sum of all layers, which is why the transfer segment in the demo above is narrow relative to the total. The cost of this overlap is a coupling between the two pools: they must agree on layer boundaries, paging layout and per-layer chunk sizes, which is exactly the interface NIXL and the vLLM/SGLang connectors standardise. Decode's first-token latency is then bounded by the transfer tail plus local attention, not by the full cache size.
When disaggregation loses
Short prompts and idle pools
Disaggregation is not free, and for the wrong workload it is strictly worse. It loses when:
- Prompts are short. The fixed setup cost (connection, memory registration, handshake) plus the transfer latency exceeds the prefill tax you avoid. The demo above computes the break-even length; for a 50 GB/s link and GQA it is a few hundred tokens.
- The pools are small or bursty. Two pools of 4 GPUs each cannot absorb each other's idle time. Disaggregation needs enough replicas that a busy period in one pool overlaps an idle period in the other.
- Traffic is low. Under light load, colocation's interference is invisible and splitting only adds a copy and a second failure domain.
- Bandwidth is scarce. If the KV transfer competes with the decode workers' own traffic, the copy shows up as ITL jitter.
KV-aware routing across pools
Where does the cache live?
Once decode workers are stateful caches, the router cannot be round-robin. It must prefer the decode worker that already holds (or can cheaply fetch) the request's prefix, balancing cache affinity against load and transfer cost — the same problem Part 8 describes for a single pool, now with two hops. A production router scores candidates on prefix match length, current batch occupancy, and estimated transfer time, and it must handle the case where the best-affinity worker is saturated: recomputing a few thousand cached tokens can beat a 30 ms transfer plus a queue. This is where the scheduler, the transfer engine and the cache hierarchy meet.
DeepSeek's production deployment, scaled to you
The disclosed numbers
DeepSeek's published inference overview describes one deployment unit as 32 GPUs for prefill with EP32 and 144 GPUs for decode with EP144, serving all V3/R1 traffic over 24 hours: 608 billion input tokens and 168 billion output tokens. 342B of the input (56.3%) hit the on-disk KV cache; average decode speed was 20–22 tokens/s. The unit sizes are not arbitrary — they are the throughput-balanced split this page computes, with a lot of spare decode batch to hit the latency target. Scale your own traffic as a fraction of theirs and watch the granularity bite.
One square = one 32-GPU prefill unit or one 144-GPU decode unit. Units are not perfectly divisible.
Two autoscalers, two failure domains
The operational bill
Splitting the pools splits the operations. Prefill and decode scale on different signals — prefill on queued prefill tokens and TTFT SLO, decode on batch occupancy and ITL SLO — so they need separate autoscalers, and each can cold-start independently (a 40–90 s weight load, Part 19). A prefill surge can leave decode idle; a decode surge can queue prefilled caches waiting for a worker. And because the KV transfer is the seam, it becomes a new failure domain: a fabric hiccup turns into dropped or delayed caches, and the router must decide between retrying, recomputing, or failing over to a colocated fallback path. The reward for all of that is a system where each job runs on the right hardware with the right batch shape, and where a latency SLO is a pool-sizing parameter rather than a hope.
Cheat sheet
| Concept | What to remember |
|---|---|
| Why disaggregate | Prefill bursts stretch decode's ITL; separate pools let each phase batch optimally |
| Optimal split | \(N_p/(N_p+N_d) = r/(r+R_p/R_d)\); DeepSeek's 3.62 ratio predicts ~19% prefill |
| Transfer volume | GQA 70B ≈ 0.33 MB/token; MLA DeepSeek ≈ 0.07 MB/token (~4.7× smaller) |
| Transfer time | 1.3 GB prompt cache: ~26 ms at 50 GB/s, ~105 ms at 12.5 GB/s |
| Overlap | Stream KV per layer; critical path is the last layer, not the sum |
| Break-even | Short prompts and small/bursty pools lose; long, steady, well-connected traffic wins |
| DeepSeek unit | 32 prefill GPUs (EP32) + 144 decode GPUs (EP144); 608B in / 168B out per day |
Further reading
- Zhong et al., "DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving", OSDI 2024.
- Patel et al., "Splitwise: Efficient Generative LLM Inference Using Phase Splitting" (2024).
- Qin et al., "Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving" (2024).
- DeepSeek-AI, "DeepSeek-V3/R1 Inference System Overview" (2025-02-28).
- vLLM, disaggregated prefill/decode connectors.