Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

One base, many adapters

A fine-tune that is 1% of the model

Low-rank adaptation freezes the base weights and learns two small matrices per targeted layer, so a fine-tune becomes a small delta rather than a full model copy. That mechanism is the training guide's story (fine-tuning); the serving consequence is the interesting one. Because the adapter is so small, one base model plus N adapters is far cheaper than N models, and the adapter can be swapped per request. A rank-16 adapter on a 70B model is roughly 150–200 MB in BF16 — under 0.3% of the base — and a host can hold thousands of them in host DRAM while the base weights stay pinned in HBM. The whole game is deciding which adapters are GPU-resident at any instant and how to run many of them in one step.

💡 The shift in framing: the base model is a shared, read-mostly substrate; adapters are the tenant-specific state. Serving becomes a memory-management problem over adapter bytes, exactly the way the KV cache was a memory-management problem over activation state.
2

Batched LoRA and the SGMV kernel

One weight read, many adapters

The naive way to serve a batch that touches four adapters is four forwards, one per adapter. Each forward streams the entire base model through HBM to produce a few tokens, so the dominant cost — reading the weights — is paid four times for one scheduling opportunity. Punica's answer is Segmented Gather Matrix-Vector multiplication (SGMV): concatenate every sequence's hidden states, tag each row with its adapter, and run one grouped kernel per layer that gathers the right adapter's low-rank factors per row and applies the update in registers. The base weights are read once; the adapter factors are tiny and effectively free. The reported gain is up to 12× throughput over one-adapter-per-request serving, and it grows with the number of adapters because that is the number of weight reads you stop paying for.

$$y_i = W_0 x_i + \frac{\alpha}{r}\, B_{a(i)} A_{a(i)} x_i,\qquad a(i)=\text{adapter of row } i$$

The catch is that a grouped kernel only helps when the work is grouped. Requests for different adapters have to be admitted into the same decode step — which is a scheduler decision, not a kernel one, and it is why adapter-aware batching and fair queueing (below) are the same problem seen from two sides.

Naive serial forwards above; one fused SGMV step below. Press play.

3

S-LoRA unified paging

Adapter weights and KV in one pool

If adapters are swapped per request, GPU HBM must hold a working set of them. S-LoRA's move is to page adapter weights into the same unified memory pool that already pages the KV cache, with the same block abstraction and an LRU policy that evicts cold adapters to host DRAM. Because both live in one pool, the engine can trade adapter residency against KV residency instead of reserving a fixed slice for each. S-LoRA reports serving thousands of concurrent adapters on a single GPU, with tensor-parallel adapter shards so the paging cost is split across the same GPUs that run the base model. The residency policy is a pure cache problem, and it has a cache's behaviour: a Zipf popularity skew means a tiny resident set serves most of the traffic.

Hit rate versus GPU-resident adapter capacity, for the current popularity skew.

4

Adapter cold-start

Milliseconds, not the minutes a base model costs

A base-model cold start is dominated by pulling tens of gigabytes of weights and building CUDA graphs, which is where the 40–90 s figure from Part 19 comes from. An adapter is three orders of magnitude smaller, so the arithmetic is friendly: 160 MB over PCIe Gen4 x16 at ~25 GB/s is about 6.4 ms, and over Gen5 at ~64 GB/s about 2.5 ms. The costs that remain are the ones people forget — a graph capture or kernel-shape selection for a rank/module combination not seen before, and the fact that the first request to a cold adapter pays the whole fetch inside its own TTFT. That is why a miss shows up as a bimodal TTFT distribution rather than a uniform slowdown, and why the cache hit rate from the demo above is a latency SLO input, not just a bandwidth number.

💡 Design rule: keep the adapter catalog on a fast local tier and pre-warm the top of the Zipf head at startup. Paging catches the tail; pre-warming keeps the p99 of the head, which is where the latency budget actually lives.
5

RPM, TPM and concurrency limits

Three dials that are not interchangeable

A quota surface exposes three independent knobs, and they fail in different ways. Requests per minute (RPM) caps call count and is trivially gamed by cheap requests — a tenant can spend its whole RPM budget on one-token prompts and use almost no GPU. Tokens per minute (TPM) caps the actual work and is the closest thing to a cost proxy, but a single 128K-token prompt can consume a minute's budget in one call. Concurrency caps in-flight sequences, which is the real constraint on KV residency and decode batching; it does not care about token count at all. A serious gateway limits all three, because each one is the failure mode of the other two.

RPM

Caps call count. Cheap-request flooding spends it instantly; a long prompt costs one unit. Enforce at the gateway before routing.

TPM

Caps work. Needs streaming-aware accounting: a request's token total is unknown until it finishes, so reserve against the budget and reconcile on completion.

Concurrency

Caps resident sequences and therefore KV pressure. The one dial that protects the engine when two tenants both go long-context at once.

SLO per tenant

The point of the other three. Quotas are admission control; the SLO is what you promised once admitted, and it needs its own isolation story.

6

Token buckets versus sliding windows

Burst, sustained rate, and the 429 tail

A token bucket is a bucket of capacity B tokens that refills at r tokens per second; each request drains its cost, and a request that cannot be paid for gets a 429. Its virtue is that it tolerates burst up to B while enforcing an average of r, which matches how real workloads arrive: a batch job flush, then a quiet period. A fixed window does the opposite — it allows 2× the limit across a window boundary, because the last second of one window and the first of the next are both measured independently. A sliding window logs every request in the trailing interval, which is precise but costs memory proportional to request rate and still has no notion of burst. In practice: a token bucket with a modest B for burst and a smoothing rate near the sustained SLO is the right default, and the tail behaviour after a burst is what tenants actually complain about.

A fixed 12-second trace, replayed. Change any slider to recompute it.

7

Virtual-token-counter fair queueing

Round-robin over requests is not fairness

Serve two tenants round-robin, one request at a time. If tenant S sends requests of 5 tokens and tenant L sends requests of 500 tokens, request-level round-robin gives S half the requests but roughly 1% of the tokens — and therefore roughly 1% of the GPU. Token counts are what the GPU spends, so token share is the fair currency. The Virtual Token Counter (VTC) scheduler makes that explicit: every tenant carries a counter of the tokens it has been served, and each decode step admits the requests of the tenant with the smallest counter, incrementing it by the tokens served. A tenant that arrives with a burst of long requests simply runs up its own counter and then waits behind lighter tenants, instead of monopolising the step. The reported effect (Sheng et al., OSDI 2024) is up to 2.9× better fairness and 6.3× lower worst-case latency than request-level round-robin.

Cumulative tokens served, per tenant, over the same GPU-seconds.

💡 Why it composes with LoRA: the same scheduler that equalises token share also decides which adapters share a step. VTC keeps the mix balanced; grouped batching makes the balanced mix cheap. Neither works well alone.
8

Isolation levels

From shared everything to dedicated silicon

Noisy-neighbour protection is a ladder, and each rung buys a stronger guarantee at a higher cost. Shared engine with quotas is cheapest and gives throughput isolation only if the scheduler is fair — latency is still coupled through the batch. Per-tenant queues with priorities decouple admission, but a preemption or a KV squeeze still crosses tenants. Dedicated workers per tenant (or per tenant tier) give true latency isolation and a bounded blast radius, at the cost of idle capacity and a pool per tenant. At the far end, MIG or separate GPUs give hardware isolation. The right rung is a function of the tenant's SLO and how much the provider is willing to pay for it; most production systems run two tiers — shared-with-VTC for the long tail, dedicated for the handful of tenants whose p99 is contractual.

isolation ∝ dedicated_resources · cost · lost_utilisation
9

The noisy-neighbour incident

What it looks like, and what you need to have measured

Tenant X ships a feature that fans out long-context reasoning calls at 5 p.m. on a shared pool. Three things happen at once. KV occupancy climbs until the engine starts preempting, so recompute work is injected into a batch that was already full. Adapter thrashing follows because the freed blocks are contested between KV and adapter weights in the unified pool, and every miss adds a PCIe fetch to some tenant's TTFT. And because the requests are long, they occupy decode slots for minutes, so the scheduler's token counters — if it had any — would have flagged the imbalance a minute before the latency did. Tenant Y, running short chat turns, sees p99 TTFT go from 300 ms to 6 s and starts getting 429s from its own exhausted bucket.

None of that is diagnosable from aggregate averages. The postmortem needs, per tenant and per minute: offered and admitted RPM/TPM, queue wait and KV wait separately, resident sequences, KV bytes held, adapter hit rate and PCIe swap bytes/s, preemption events, and goodput against each tenant's SLO. With those, the incident reads as a single curve — X's token share — and the fix is a policy change (a concurrency cap, a VTC guarantee, a dedicated decode pool), not a capacity purchase. Without them, it reads as "the pool got slow", and the next one reads the same way.

⚠ The trap: aggregate GPU utilisation stays near 100% through the whole incident. The system looked maximally efficient while a contractual SLO was burning. Utilisation is not goodput — the point Part 3 makes, and the reason per-tenant SLOs are the only signal that catches this class of failure.
✓

Cheat sheet

Multi-tenancy in one table

LeverWhat it buysWhat it costs / breaks
Grouped LoRA (SGMV)One base-weight read for many adapters; up to ~12× vs per-adapter forwardsNeeds the scheduler to co-batch adapters; grouped-kernel overhead at tiny batch
Unified adapter/KV pagingThousands of adapters on one GPU; residency follows Zipf, not reservationMisses add PCIe fetch to TTFT; paging competes with KV under pressure
RPM + TPM + concurrencyThree independent failure modes each pluggedAccounting a stream's tokens requires reserve-then-reconcile
Token bucketBurst up to B, average r, cheap to run429 tail after a burst; fixed windows leak 2× at boundaries
VTC fair queueingEqual token share; ~2.9× fairness, ~6.3× worst-case latencyStrong tenant guarantees need dedicated workers or MIG
Per-tenant goodputDetects the incident aggregates hideRequires per-tenant metrics at gateway and engine
📚

Further reading

References

?

Check your understanding

0/5 answered