Kubernetes, autoscaling, cold starts and routing
Everything so far described one engine on one GPU. Production is a fleet: dozens to thousands of replicas behind a gateway, a router that decides which replica answers, an autoscaler that decides how many exist, and a scheduler that decides where they run. This part is the control plane at fleet altitude — why the load balancer must know what is in each replica's KV cache, why GPU utilisation is the worst autoscaling signal available, what a cold start is actually made of, and what "scale to zero" honestly costs.
The control plane: gateway, router, replicas, autoscaler
Four components between the user and the GPU
The training guide's inference chapter treats a request as prefill then decode and stops at the engine's edge. That engine is a process on one machine; everything this part covers is the machinery that runs many of those processes as one service. Four components sit on the path, and each owns one decision. The gateway authenticates, enforces quotas (the token buckets and per-tenant limits from Part 17), terminates TLS, proxies SSE streams, and emits the request-level metrics that make the rest debuggable. The router picks an endpoint. The replicas are the GPU workers themselves, each a complete engine — often a tensor-parallel group of several GPUs that must be scheduled together. The autoscaler decides how many replicas should exist right now.
The split matters because the router and the autoscaler are answering different questions with different signals. The router sees per-request state (prompt length, prefix hash, target SLO) and must make a decision in microseconds on the hot path. The autoscaler sees aggregate telemetry over tens of seconds and makes a decision that takes a cold start — tens of seconds — to take effect. Conflating them produces the classic failure: a "least connections" router that fills every replica's KV cache with different prefixes, so no two requests ever share a cached prompt and the cache is dead weight. The gateway-plus-router split is now standardised: llm-d layers a Gateway API Inference Extension and an Endpoint Picker (EPP) on top of Envoy, NVIDIA Dynamo ships a frontend plus a KV-aware router and a planner, and KServe wraps all of it as an InferenceService with a Knative/HPA autoscaler underneath.
GPU workloads on Kubernetes
All-or-nothing resources and gang scheduling
Kubernetes was built for stateless, fungible, CPU-sized containers; a GPU worker is none of those. GPUs are exposed by the NVIDIA device plugin (or the newer Dynamic Resource Allocation API) as integer resources on a node, and the default request is a whole device. There is no fractional GPU unless you explicitly carve MIG slices, and a tensor-parallel replica spanning four GPUs has to land on four GPUs that can talk over NVLink — which for a GB200 NVL72 domain means the same rack, an NVLink fabric of roughly 130 TB/s of aggregate bisection. A single unschedulable rank dooms the whole replica, so inference deployments use gang scheduling (Kueue, Volcano, or the scheduler's own coscheduling) rather than hoping four independent pods co-locate.
The practical consequences are worth stating plainly. Scaling granularity is the TP degree times the data-parallel degree — a TP=8 model moves in units of eight GPUs, so "add one replica" is really "add eight GPUs", and the autoscaler's minimum step is much coarser than for a web service. Node pools are typically homogeneous and tainted so that a 70B FP8 deployment does not accidentally land on a pool built for a 7B. And the model itself is 140 GB to several hundred GB, so pulling it onto a node is a real, slow, occasionally-throttled operation — which is the entire subject of the cold-start section below. The three stack options map cleanly onto this: llm-d is the Kubernetes-native, Gateway-API-based composition of vLLM workers with an EPP router; NVIDIA Dynamo is a higher-level serving framework with a KV router, a planner and NIXL-based KV transfer; and KServe is the general model-serving control plane that gives you canaries, revisioned endpoints and scale-to-zero in exchange for a more opinionated deployment shape.
The balancer must be KV-aware
Round-robin is a cache flush disguised as fairness
Part 8 established prefix caching: an agent's system prompt, a shared document, or a tool schema is the same token sequence on every call, so a cache hit turns a full prefill into a shorter one — 2,100 ms down to 190 ms in the canonical case. At fleet altitude, an ordinary least-connections or round-robin balancer destroys that property, because the replica that answers your next request is a different replica than the one that cached your prefix. The cache is per-replica state; a router that ignores it guarantees a cache miss. The fix is a router that scores endpoints on cache affinity first and load second: hash the prefix, find the replicas that hold it, and among those pick the least loaded. llm-d's Endpoint Picker and Dynamo's KV router both do exactly this, and consistent hashing is the usual way to keep a tenant's prefixes pinned to a stable replica even as the pool resizes.
The payoff is large and measurable. In vLLM's Mooncake integration on agentic coding traces, cache-aware routing took the hit rate from 1.7% to 92.2%, for 3.8× higher throughput and 46× lower p50 TTFT. The cost is a new failure mode: affinity fights balance, and a hot tenant can pile onto one replica while its neighbours idle. Every serious implementation therefore blends the two scores and has a fallback — if the affined replica is saturated, take the miss and send the request elsewhere. The demo below puts four policies through the same skewed tenant traffic and reports hit rate and tail TTFT.
Four routing policies over 4,000 requests from 32 tenants with a shared system prefix. Hit rate above, TTFT below.
Autoscaling signals: GPU utilisation is a lie
The signal that arrives 90 seconds too late
The obvious autoscaling signal is GPU utilisation, and it is the worst one available. An LLM decode step is memory-bandwidth-bound: it streams the weights through HBM and does very little arithmetic per byte (Part 1's ~1–2 FLOPs/byte), so SM occupancy can sit near 30% while the GPU is completely saturated on bandwidth and requests are queueing. Utilisation is also measured after the fact — the agent samples it on a scrape interval, the value is averaged over the window, and by the time it climbs the queue has already formed and the first cold-started replica is still 40 s from being ready. A utilisation-based autoscaler is structurally a lagging indicator pointed at a resource that does not measure the bottleneck.
Two signals do work. Queue depth — requests waiting to be admitted, and KV blocks waiting to be freed — is a leading indicator: it rises the instant offered load exceeds capacity and falls as capacity catches up. p99 TTFT is a lagging indicator but it is the truthful one: it measures what the user feels, and it is the only signal that captures the cases where the bottleneck is scheduling, preemption, or a slow prefix cache rather than raw GPU throughput. In practice you drive the autoscaler from queue depth (or a tokens-in-flight proxy) and guard the down-scale with an SLO-based signal, with a cooldown that is at least as long as the cold start. The demo makes the difference visible: switch the signal and watch the same traffic produce very different replica curves and SLO violations.
Ten minutes of traffic, 2 s steps. Top: offered load against ready capacity. Bottom: p99 TTFT (left axis) and replica count (right axis).
Cold start, broken down
40–90 seconds, and the term that survives every fix
A cold start is not one delay; it is a stack of them, and each lever removes a named term. The table below decomposes a ~92 s baseline: pulling a multi-gigabyte CUDA-plus-PyTorch image, initialising the process and CUDA context, loading tens of gigabytes of weights from disk, capturing CUDA graphs, warming the engine, JIT-compiling or autotuning kernels, and finally the engine's own config load. The levers are well understood. A slim image (a runtime-only, pre-baked image instead of a base with the full toolkit) cuts the pull from ~40 s to single digits. Local weights on instance-store NVMe or a node-local cache cut disk load from network-FS speeds. NVIDIA Run:ai Model Streamer goes further by reading the safetensors concurrently over many threads and streaming layers to the GPU as they arrive: documented 43.7 s down to 7.5 s for a 15 GB Llama 3 8B on IO2 SSD, a ~5.8× gain. Graph reuse shares pre-captured CUDA graphs across replicas so capture happens once per shape, not once per process. And a snapshot (CRIU- or CUDA-checkpoint-based) restores an already-initialised process, collapsing process init, graph capture, warmup and autotune into a fast memory restore.
Toggle the levers and watch the residual. With everything on, the cold start lands near 21 s — and the dominant remaining term is still the weight load. That is the honest conclusion: once images are slim and the process is snapshotted, the model bytes are the floor, so the last levers are storage bandwidth and concurrency, plus a warm pool if even 21 s is too slow.
Each toggle removes a stage; the bar eases to the new total. Baseline 92 s.
Honest scale-to-zero
The break-even gap is a latency decision, not a dollar one
Scaling to zero is easy to justify on GPU dollars and hard to justify on latency. Pure GPU-seconds always favour scale-to-zero: a cold start bills 45 s of one GPU, while keeping a replica warm for an idle hour bills 3,600. The reason to keep a replica warm is that the first request after a scale-to-zero pays the entire cold start inside its own TTFT — a 40 s p99 is not a p99, it is an outage with good average latency hiding it. The correct framing is a break-even idle gap: if the expected gap to the next request is shorter than roughly the cold start plus the value of an SLO breach expressed in GPU-seconds, keep the replica warm. With a 2 USD breach cost and 3 USD/hour GPUs that gap is about 41 minutes; below it, warming is cheaper than repeatedly cold-starting into a timeout. The demo shows both the pure-dollar saving and the SLO-adjusted break-even, because the second is the one that determines whether scale-to-zero is defensible. When it truly is — overnight, a low-traffic region, an internal tool — snapshots and a small warm floor (a minimum replica count of one, or a warm pool the Knative/Dynamo planner draws from) are what make the first request survivable.
Hourly cost of one warm replica versus the SLO-adjusted cost of cold starts, across arrival rate. Shaded: warm is cheaper.
Spot preemption and draining
What you do with the last two minutes
Spot and preemptible GPUs are 50–80% cheaper and can be reclaimed at any time, usually with a short notice window (a two-minute warning on AWS, 30 s on some other clouds). What happens during that window decides how expensive the discount really is. Three policies span the space. Hard kill terminates immediately: every in-flight request is lost, the replica restarts cold, and clients retry. Finish in flight stops admitting new work and drains the running requests before terminating — no lost work, but the node is unavailable for new arrivals for the duration of its longest remaining generation, and its replacement only starts after the drain. Migrate KV streams the in-flight KV caches to a surviving replica over the fabric (the same NIXL/transfer path as Part 13) and continues there, so nothing is lost and the drain is a transfer instead of a wait. Each policy implies a minimum spot discount for it to beat on-demand: the demo computes that break-even discount from the reclaim rate, the work at risk, and the cost of the interruption. The point is that a 15% discount can be a bad trade for hard-kill fleets and a good one for migrating fleets, and the difference is policy, not price.
Required spot discount per drain policy against the discount actually offered. A bar to the right of the marker is not worth it.
Multi-region and the cache you lose
Failover is a cold start plus a cache-miss storm
A prefix cache is local state, and it does not replicate. That single fact makes multi-region routing harder than multi-region routing for stateless web services. In steady state the right objective is not geographic proximity but cache and residency: route a user to the region (and within it, the replica) that already holds their prefix, because a warm cache at higher RTT is faster than a cold prefill nearby. But the moment a region fails, every prefix cached there is gone, and the surviving region absorbs the traffic with an empty cache — a thundering herd of recomputation exactly when the system is least able to afford it. A global load balancer with health checks gives you availability; it does not give you the cache back. What helps: per-region warm pools sized for failover, sticky routing keyed on prefix hash so a tenant does not bounce, a shared KV tier (LMCache-style) that survives pod restarts, and pre-warming the top prefixes in a region before shifting traffic to it. This is Part 8's cache economics applied at fleet scale, and it is why "add a region" is a capacity plan, not just a DNS change. Treat it as capacity planning for the next part — the cost and capacity model is where these decisions get a dollar figure, and the glossary is where every term here is indexed.
Cheat sheet
The control plane in one table
| Component / lever | What it buys | What it costs / breaks |
|---|---|---|
| Gateway | Auth, quotas, streaming proxy, per-request metrics | One more hot-path hop; quota accounting must reserve-then-reconcile for streams |
| KV-aware router | 2,100 → 190 ms TTFT on a hit; 1.7% → 92.2% fleet hit rate (vLLM×Mooncake) | Affinity fights balance; needs a fallback when the affined replica is hot |
| Replicas = whole TP groups | NVLink-domain locality; one sharding decision | Scaling granularity is TP×DP GPUs; one bad rank dooms the replica (gang schedule it) |
| Autoscale signal: GPU util | Nothing useful — memory-bound decode keeps util low while latency spikes | Lags by the scrape window; reacts ~90 s late even on a clean step change |
| Autoscale signal: queue depth | Leading indicator; rises the instant load exceeds capacity | Needs a cooldown ≥ cold start or it flaps and provisions into a falling queue |
| Autoscale signal: p99 TTFT | The truthful user-facing signal; catches scheduling and cache failures | Lagging — it is already bad before it moves; guard down-scale with it |
| Slim image + local weights | Cuts pull and disk load, the two biggest cold-start terms | Image bake pipeline and node-local storage to keep fresh |
| Model Streamer / graph reuse / snapshot | 43.7 → 7.5 s weight load; capture once; restore pre-warmed process | Concurrency for streaming; graph/engine coupling; snapshot tooling and restore correctness |
| Scale-to-zero | Saves idle GPU dollars | First request pays the whole cold start; break-even gap ≈ cold start + breach value / price |
| Spot + migrate KV | Discount with no lost work; drain becomes a transfer | Needs a spare replica and fabric bandwidth; the other policies need a real discount |
Further reading
References
- llm-d — Kubernetes-native distributed inference: Gateway API Inference Extension, the Endpoint Picker (EPP) and KV-aware scoring, with vLLM workers.
- NVIDIA Dynamo — disaggregated serving with a KV-aware router, a planner for autoscaling, and NIXL for KV transfer; see also NIXL.
- KServe —
InferenceService, revisioned endpoints, canaries, and Knative/HPA autoscaling including scale-to-zero. - NVIDIA Run:ai, Model Streamer and the NVIDIA developer blog, Reducing Cold Start Latency for LLM Inference (2025) — 43.7 s → 7.5 s for a 15 GB model.
- Zheng et al., SGLang / RadixAttention and the vLLM × Mooncake cache-aware routing results — the fleet-altitude case for prefix affinity, introduced in Part 8.
- Kubernetes docs, Scheduling, Preemption and Eviction, and Kueue's gang-scheduling for co-located GPU workers.