LoRA, quotas and noisy neighbours
A served model is rarely one model. Teams fine-tune it into hundreds of adapters, tenants buy different rate limits, and one customer's 200K-token reasoning burst can wreck everyone else's p99. This part is about the machinery that keeps many tenants on one base: grouped LoRA kernels and adapter paging underneath, quota and fairness policy on top, and the measurements that tell you which of the two is actually failing.
One base, many adapters
A fine-tune that is 1% of the model
Low-rank adaptation freezes the base weights and learns two small matrices per targeted layer, so a fine-tune becomes a small delta rather than a full model copy. That mechanism is the training guide's story (fine-tuning); the serving consequence is the interesting one. Because the adapter is so small, one base model plus N adapters is far cheaper than N models, and the adapter can be swapped per request. A rank-16 adapter on a 70B model is roughly 150–200 MB in BF16 — under 0.3% of the base — and a host can hold thousands of them in host DRAM while the base weights stay pinned in HBM. The whole game is deciding which adapters are GPU-resident at any instant and how to run many of them in one step.
Batched LoRA and the SGMV kernel
One weight read, many adapters
The naive way to serve a batch that touches four adapters is four forwards, one per adapter. Each forward streams the entire base model through HBM to produce a few tokens, so the dominant cost — reading the weights — is paid four times for one scheduling opportunity. Punica's answer is Segmented Gather Matrix-Vector multiplication (SGMV): concatenate every sequence's hidden states, tag each row with its adapter, and run one grouped kernel per layer that gathers the right adapter's low-rank factors per row and applies the update in registers. The base weights are read once; the adapter factors are tiny and effectively free. The reported gain is up to 12× throughput over one-adapter-per-request serving, and it grows with the number of adapters because that is the number of weight reads you stop paying for.
The catch is that a grouped kernel only helps when the work is grouped. Requests for different adapters have to be admitted into the same decode step — which is a scheduler decision, not a kernel one, and it is why adapter-aware batching and fair queueing (below) are the same problem seen from two sides.
Naive serial forwards above; one fused SGMV step below. Press play.
S-LoRA unified paging
Adapter weights and KV in one pool
If adapters are swapped per request, GPU HBM must hold a working set of them. S-LoRA's move is to page adapter weights into the same unified memory pool that already pages the KV cache, with the same block abstraction and an LRU policy that evicts cold adapters to host DRAM. Because both live in one pool, the engine can trade adapter residency against KV residency instead of reserving a fixed slice for each. S-LoRA reports serving thousands of concurrent adapters on a single GPU, with tensor-parallel adapter shards so the paging cost is split across the same GPUs that run the base model. The residency policy is a pure cache problem, and it has a cache's behaviour: a Zipf popularity skew means a tiny resident set serves most of the traffic.
Hit rate versus GPU-resident adapter capacity, for the current popularity skew.
Adapter cold-start
Milliseconds, not the minutes a base model costs
A base-model cold start is dominated by pulling tens of gigabytes of weights and building CUDA graphs, which is where the 40–90 s figure from Part 19 comes from. An adapter is three orders of magnitude smaller, so the arithmetic is friendly: 160 MB over PCIe Gen4 x16 at ~25 GB/s is about 6.4 ms, and over Gen5 at ~64 GB/s about 2.5 ms. The costs that remain are the ones people forget — a graph capture or kernel-shape selection for a rank/module combination not seen before, and the fact that the first request to a cold adapter pays the whole fetch inside its own TTFT. That is why a miss shows up as a bimodal TTFT distribution rather than a uniform slowdown, and why the cache hit rate from the demo above is a latency SLO input, not just a bandwidth number.
RPM, TPM and concurrency limits
Three dials that are not interchangeable
A quota surface exposes three independent knobs, and they fail in different ways. Requests per minute (RPM) caps call count and is trivially gamed by cheap requests — a tenant can spend its whole RPM budget on one-token prompts and use almost no GPU. Tokens per minute (TPM) caps the actual work and is the closest thing to a cost proxy, but a single 128K-token prompt can consume a minute's budget in one call. Concurrency caps in-flight sequences, which is the real constraint on KV residency and decode batching; it does not care about token count at all. A serious gateway limits all three, because each one is the failure mode of the other two.
RPM
Caps call count. Cheap-request flooding spends it instantly; a long prompt costs one unit. Enforce at the gateway before routing.
TPM
Caps work. Needs streaming-aware accounting: a request's token total is unknown until it finishes, so reserve against the budget and reconcile on completion.
Concurrency
Caps resident sequences and therefore KV pressure. The one dial that protects the engine when two tenants both go long-context at once.
SLO per tenant
The point of the other three. Quotas are admission control; the SLO is what you promised once admitted, and it needs its own isolation story.
Token buckets versus sliding windows
Burst, sustained rate, and the 429 tail
A token bucket is a bucket of capacity B tokens that refills at r tokens per second; each request drains its cost, and a request that cannot be paid for gets a 429. Its virtue is that it tolerates burst up to B while enforcing an average of r, which matches how real workloads arrive: a batch job flush, then a quiet period. A fixed window does the opposite — it allows 2× the limit across a window boundary, because the last second of one window and the first of the next are both measured independently. A sliding window logs every request in the trailing interval, which is precise but costs memory proportional to request rate and still has no notion of burst. In practice: a token bucket with a modest B for burst and a smoothing rate near the sustained SLO is the right default, and the tail behaviour after a burst is what tenants actually complain about.
A fixed 12-second trace, replayed. Change any slider to recompute it.
Virtual-token-counter fair queueing
Round-robin over requests is not fairness
Serve two tenants round-robin, one request at a time. If tenant S sends requests of 5 tokens and tenant L sends requests of 500 tokens, request-level round-robin gives S half the requests but roughly 1% of the tokens — and therefore roughly 1% of the GPU. Token counts are what the GPU spends, so token share is the fair currency. The Virtual Token Counter (VTC) scheduler makes that explicit: every tenant carries a counter of the tokens it has been served, and each decode step admits the requests of the tenant with the smallest counter, incrementing it by the tokens served. A tenant that arrives with a burst of long requests simply runs up its own counter and then waits behind lighter tenants, instead of monopolising the step. The reported effect (Sheng et al., OSDI 2024) is up to 2.9× better fairness and 6.3× lower worst-case latency than request-level round-robin.
Cumulative tokens served, per tenant, over the same GPU-seconds.
Isolation levels
From shared everything to dedicated silicon
Noisy-neighbour protection is a ladder, and each rung buys a stronger guarantee at a higher cost. Shared engine with quotas is cheapest and gives throughput isolation only if the scheduler is fair — latency is still coupled through the batch. Per-tenant queues with priorities decouple admission, but a preemption or a KV squeeze still crosses tenants. Dedicated workers per tenant (or per tenant tier) give true latency isolation and a bounded blast radius, at the cost of idle capacity and a pool per tenant. At the far end, MIG or separate GPUs give hardware isolation. The right rung is a function of the tenant's SLO and how much the provider is willing to pay for it; most production systems run two tiers — shared-with-VTC for the long tail, dedicated for the handful of tenants whose p99 is contractual.
The noisy-neighbour incident
What it looks like, and what you need to have measured
Tenant X ships a feature that fans out long-context reasoning calls at 5 p.m. on a shared pool. Three things happen at once. KV occupancy climbs until the engine starts preempting, so recompute work is injected into a batch that was already full. Adapter thrashing follows because the freed blocks are contested between KV and adapter weights in the unified pool, and every miss adds a PCIe fetch to some tenant's TTFT. And because the requests are long, they occupy decode slots for minutes, so the scheduler's token counters — if it had any — would have flagged the imbalance a minute before the latency did. Tenant Y, running short chat turns, sees p99 TTFT go from 300 ms to 6 s and starts getting 429s from its own exhausted bucket.
None of that is diagnosable from aggregate averages. The postmortem needs, per tenant and per minute: offered and admitted RPM/TPM, queue wait and KV wait separately, resident sequences, KV bytes held, adapter hit rate and PCIe swap bytes/s, preemption events, and goodput against each tenant's SLO. With those, the incident reads as a single curve — X's token share — and the fix is a policy change (a concurrency cap, a VTC guarantee, a dedicated decode pool), not a capacity purchase. Without them, it reads as "the pool got slow", and the next one reads the same way.
Cheat sheet
Multi-tenancy in one table
| Lever | What it buys | What it costs / breaks |
|---|---|---|
| Grouped LoRA (SGMV) | One base-weight read for many adapters; up to ~12× vs per-adapter forwards | Needs the scheduler to co-batch adapters; grouped-kernel overhead at tiny batch |
| Unified adapter/KV paging | Thousands of adapters on one GPU; residency follows Zipf, not reservation | Misses add PCIe fetch to TTFT; paging competes with KV under pressure |
| RPM + TPM + concurrency | Three independent failure modes each plugged | Accounting a stream's tokens requires reserve-then-reconcile |
| Token bucket | Burst up to B, average r, cheap to run | 429 tail after a burst; fixed windows leak 2× at boundaries |
| VTC fair queueing | Equal token share; ~2.9× fairness, ~6.3× worst-case latency | Strong tenant guarantees need dedicated workers or MIG |
| Per-tenant goodput | Detects the incident aggregates hide | Requires per-tenant metrics at gateway and engine |
Further reading
References
- Hu et al., LoRA: Low-Rank Adaptation of Large Language Models, ICLR 2022.
- Chen et al., Punica: Multi-Tenant LoRA Serving, MLSys 2024 — SGMV and the grouped-kernel design.
- Sheng et al., S-LoRA: Serving Thousands of Concurrent LoRA Adapters — unified paging of adapters and KV.
- Sheng et al., Fairness in Serving Large Language Models, OSDI 2024 — the Virtual Token Counter.
- The paged memory and scheduler parts for the block machinery this page builds on.