Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

One pool, two jobs

The ITL spike you can see

The training guide's inference chapter framed prefill as compute-bound and decode as bandwidth-bound. On a shared engine, those two workloads contend for the same SMs: while a 4,000-token prompt is prefilled, the running decode batch waits its turn, and a user watching tokens stream sees a stall. Part 3 named the metric — inter-token latency, or ITL — and this is the mechanism that makes its p99 ugly even when throughput looks healthy.

Chunked prefill (Part 6) mitigates this by slicing prompts into iteration-sized pieces, but the total work is unchanged: some iterations still carry prefill, and the decode step is longer. Disaggregation attacks the problem by routing each request's two phases to different GPUs. Prefill workers batch many prompts together, saturating compute with large matrix multiplies; decode workers run a steady, latency-optimised loop with a fixed batch. The prefill worker sends the resulting KV cache to the decode worker over the network, and the decode worker streams tokens back. The demo below replays both designs: the top timeline is colocated and shows a prefill spike every few iterations; the bottom keeps decode flat and exposes a small transfer segment.

Each bar is one engine iteration, stacked by job. Play to advance the timeline.

2

DistServe, Splitwise and Mooncake

Three papers, one architecture

DistServe (OSDI 2024) introduced the formulation and reported serving 7.4× more requests or 12.6× tighter SLO by splitting the phases; the widely quoted 20× is a rounded composite, not a single measurement. Splitwise (Microsoft) made the same split for a different reason — phase-specific hardware, since prefill wants compute and decode wants memory capacity. Mooncake (Moonshot/Kimi) went further and made the KV cache itself the centre of the system, with a KVCache-centric scheduler and a transfer engine, reporting up to 525% throughput in simulation and 75% more real requests. The same architecture now ships as first-class support in vLLM and SGLang via disaggregation connectors.

💡 The insight that unifies them: once the KV cache can move independently, the GPU becomes a stateless compute resource and the cache becomes the durable thing. That is what makes per-phase sizing, per-phase autoscaling and cache-aware routing possible.
3

Sizing the pools

The input:output ratio decides everything

Pool sizing is an arithmetic question. Let \(R_p\) be prefill tokens/s per GPU and \(R_d\) decode tokens/s per GPU. For a workload with \(P\) input and \(O\) output tokens per request, the GPU-seconds needed per request are \(a = P/R_p\) for prefill and \(b = O/R_d\) for decode, so the throughput-balanced split is

$$ \frac{N_p}{N_p+N_d} = \frac{a}{a+b} = \frac{r}{r + R_p/R_d}\quad\text{with } r = P/O. $$

Because decode is bandwidth-bound and batched, \(R_p/R_d\) is large — with representative values of 12,000 prefill and 800 decode tokens/s per GPU it is 15. That makes the optimal prefill share \(r/(r+15)\): a chat workload at \(r=4\) wants ~21% prefill, summarisation at \(r=20\) wants ~57%, and a reasoning workload at \(r=1/6\) wants ~1%. DeepSeek's disclosed ratio of 608B input to 168B output tokens (\(r=3.62\)) predicts ~19% prefill — and their deployment runs 32 prefill to 144 decode GPUs, or 18.2%. Use the presets to see each workload, then drag the split away from optimal and watch queues form.

Bars show utilisation of each pool relative to the throughput-balanced allocation; the marker is the optimal prefill share.

4

Moving KV: NIXL, Mooncake Transfer Engine, RDMA

The cost you pay for freedom

The transfer is the whole bargain. A prefill worker holds a KV cache of \(2\,L\,n_{kv}d_{head}T\) bytes per sequence; for a 70B GQA model that is about 0.33 MB per token, so a 4,000-token prompt is ~1.3 GB. Moving that over a 400 Gb/s (50 GB/s) RDMA link takes ~26 ms; over 100 GbE it takes ~105 ms; over NVLink it is under 3 ms. Transfer engines (NVIDIA's NIXL, Mooncake Transfer Engine, LMCache) exist to drive RDMA queues at line rate, register GPU memory once and pipeline chunk-by-chunk, so the copy overlaps the decode that follows. MLA models cut the payload ~4.7× — DeepSeek's latent is 70 KB/token versus ~328 KB for a same-geometry GQA cache — which is a large part of why disaggregation is practical for them.

The crossing point is the minimum prompt length at which relocating the KV beats recomputing prefill in place.

5

Layer-wise overlap

Don't wait for the whole cache

The naive transfer waits for prefill to finish, sends the entire cache, then starts decode. Real systems pipeline at layer granularity: the prefill worker streams layer \(i\)'s KV as soon as it is computed, and the decode worker begins prefilling its own layer \(i\) state and starts generating the first token as the tail of the cache arrives. The critical path becomes the transfer of the last layer, not the sum of all layers, which is why the transfer segment in the demo above is narrow relative to the total. The cost of this overlap is a coupling between the two pools: they must agree on layer boundaries, paging layout and per-layer chunk sizes, which is exactly the interface NIXL and the vLLM/SGLang connectors standardise. Decode's first-token latency is then bounded by the transfer tail plus local attention, not by the full cache size.

6

When disaggregation loses

Short prompts and idle pools

Disaggregation is not free, and for the wrong workload it is strictly worse. It loses when:

⚠️ The honest rule: disaggregate when prompts are long, load is high and steady, and the link is fast relative to the cache. Otherwise, chunked prefill in one pool is simpler and often faster.
7

KV-aware routing across pools

Where does the cache live?

Once decode workers are stateful caches, the router cannot be round-robin. It must prefer the decode worker that already holds (or can cheaply fetch) the request's prefix, balancing cache affinity against load and transfer cost — the same problem Part 8 describes for a single pool, now with two hops. A production router scores candidates on prefix match length, current batch occupancy, and estimated transfer time, and it must handle the case where the best-affinity worker is saturated: recomputing a few thousand cached tokens can beat a 30 ms transfer plus a queue. This is where the scheduler, the transfer engine and the cache hierarchy meet.

8

DeepSeek's production deployment, scaled to you

The disclosed numbers

DeepSeek's published inference overview describes one deployment unit as 32 GPUs for prefill with EP32 and 144 GPUs for decode with EP144, serving all V3/R1 traffic over 24 hours: 608 billion input tokens and 168 billion output tokens. 342B of the input (56.3%) hit the on-disk KV cache; average decode speed was 20–22 tokens/s. The unit sizes are not arbitrary — they are the throughput-balanced split this page computes, with a lot of spare decode batch to hit the latency target. Scale your own traffic as a fraction of theirs and watch the granularity bite.

One square = one 32-GPU prefill unit or one 144-GPU decode unit. Units are not perfectly divisible.

9

Two autoscalers, two failure domains

The operational bill

Splitting the pools splits the operations. Prefill and decode scale on different signals — prefill on queued prefill tokens and TTFT SLO, decode on batch occupancy and ITL SLO — so they need separate autoscalers, and each can cold-start independently (a 40–90 s weight load, Part 19). A prefill surge can leave decode idle; a decode surge can queue prefilled caches waiting for a worker. And because the KV transfer is the seam, it becomes a new failure domain: a fabric hiccup turns into dropped or delayed caches, and the router must decide between retrying, recomputing, or failing over to a colocated fallback path. The reward for all of that is a system where each job runs on the right hardware with the right batch shape, and where a latency SLO is a pool-sizing parameter rather than a hope.

📌 Cross-links: the KV hierarchy is Part 8; the scheduler that decides admit vs wait is Part 7; the rack-scale NVLink domain that makes the transfer cheap is Part 2 and Part 12.

Cheat sheet

ConceptWhat to remember
Why disaggregatePrefill bursts stretch decode's ITL; separate pools let each phase batch optimally
Optimal split\(N_p/(N_p+N_d) = r/(r+R_p/R_d)\); DeepSeek's 3.62 ratio predicts ~19% prefill
Transfer volumeGQA 70B ≈ 0.33 MB/token; MLA DeepSeek ≈ 0.07 MB/token (~4.7× smaller)
Transfer time1.3 GB prompt cache: ~26 ms at 50 GB/s, ~105 ms at 12.5 GB/s
OverlapStream KV per layer; critical path is the last layer, not the sum
Break-evenShort prompts and small/bursty pools lose; long, steady, well-connected traffic wins
DeepSeek unit32 prefill GPUs (EP32) + 144 decode GPUs (EP144); 608B in / 168B out per day

Further reading

?

Check your understanding

0/5 answered