The machine: HBM, interconnect and the accelerator zoo
Part 1 turned serving into a bandwidth problem: a decode step reads the weights once and writes one token. This part makes that concrete. We rank the accelerators by the numbers that actually decide tokens per second, watch models fail to fit, put a price on the network a tensor-parallel step has to cross, and end with joules per token — the number that eventually shows up on the electricity bill.
HBM bandwidth is the headline number
Why spec sheets lead with TB/s, not TFLOP/s
A prefill pass multiplies a whole prompt against the weights at once, so it saturates the tensor cores and is compute-bound. A decode step multiplies one token against the same weights, so it is memory-bandwidth-bound: the arithmetic intensity is roughly one to two FLOPs per byte read, far below the ridge point of any modern accelerator. The practical decode step time is therefore set by how fast HBM can stream the weights in, not by how many FLOPs the chip can issue.
The first term barely changes with batch size — the weights are read once per step regardless of how many sequences share that step — which is exactly why batching lifts throughput without a matching latency cost. The second term grows linearly with batch, and where the two cross is the point at which a decode step flips from memory-bound to compute-bound. H100 SXM sits at 3.35 TB/s of HBM and 989 dense BF16 TFLOP/s; B200 doubles the bandwidth to 8 TB/s and the dense throughput to 2,250 BF16 TFLOP/s. The bandwidth doubling is the part that matters for decode.
The accelerator zoo
Pick up to four; switch the metric
The comparison below uses dense (non-sparse) throughput where vendors quote sparsity, because serving without 2:4 sparsity runs at the dense rate. Switch the metric and the ranking rearranges violently: NVIDIA leads on HBM bandwidth and NVLink, AMD leads on HBM capacity per dollar-tier, and the TPUs and Trainium trade peak FLOPs for cost and integration inside their own clouds. The decode estimate reads the model's active parameters once per token on each chip, so a 671B-parameter MoE with 37B active looks cheap per token — you pay for the active slice, not the whole file.
Dense figures. Bars are normalized to the largest selected value.
Select up to four accelerators
Capacity, and what it forces
Weights + KV + activations versus one HBM budget
Bandwidth decides how fast a step runs; capacity decides whether it runs at all. A single GPU must hold the sharded weights, the KV cache for every resident sequence, and a working set of activations. The weights scale with parameter count and precision, the KV cache scales with context length, concurrency and the attention scheme (Part 4 sizes it properly), and activations scale with batch and hidden width. The moment the stack crosses the HBM capacity the engine has to shard, quantize, or cut context — there is no third option.
The horizontal line is the accelerator's HBM capacity.
NVLink and the all-reduce tax
Tensor parallelism is a network decision
Tensor parallelism splits a matrix multiply across GPUs, but every layer's partial sums have to be combined before the next layer can start. That is two all-reduces per transformer layer — one after attention, one after the MLP — and each one moves the full activation tensor across the link. On a fast NVLink domain the collective hides under the compute, and the tax is invisible. Move the same shard to PCIe or Ethernet and communication dominates: the link stops hiding it, and adding GPUs can make each step slower. This is why tensor parallelism essentially never leaves NVLink, while pipeline and expert parallelism tolerate slower fabrics.
Added latency per decode step versus tensor-parallel degree.
InfiniBand, RoCE and Ethernet
The fabric between nodes
Inside a node, NVLink carries tensor-parallel collectives. Between nodes, the options are InfiniBand (NDR at 400 Gb/s ≈ 50 GB/s per port, XDR at 800 Gb/s), RoCEv2 (RDMA over Converged Ethernet, which is what most cloud fleets actually run), and plain Ethernet. Their job in serving is different from training: they carry KV transfers between disaggregated prefill and decode pools, expert-parallel all-to-all traffic for MoE, and cache-fetch traffic from a remote KV tier. Two things matter. First, a single KV transfer is a latency-sensitive burst, so RDMA (IB or RoCE) beats TCP/IP by a wide margin. Second, expert all-to-all is bandwidth-bound and scales with the rack's aggregate fabric, which is why NVLink-based racks exist at all.
Relative time to move 1 GB of KV between two nodes.
GB200 NVL72 as one big GPU
A 130 TB/s single NVLink domain
GB200 NVL72 wires 72 Blackwell GPUs into one NVLink domain: 1.8 TB/s per GPU, roughly 130 TB/s aggregate, across a liquid-cooled rack drawing on the order of 120 kW. The aggregate is what lets a 671B-parameter model run tensor-parallel and expert-parallel without ever touching Ethernet — the model is effectively one compute unit with 13.4 TB of HBM3e at 576 TB/s aggregate. For serving, that transforms the sharding question: instead of fitting a model onto one 80 GB GPU, you fit it onto a rack and treat the rack as the allocation unit. The cost is that you cannot scale to zero and you pay for the whole rack's power whether it is busy or not.
Per-GPU versus per-rack HBM and bandwidth.
Groq and Cerebras: SRAM architectures
When the weights live on-die
The memory wall exists because HBM is off-chip. A different answer is to put the weights in on-die SRAM and give up on scaling a single model: the Groq LPU keeps on the order of 230 MB of SRAM per chip with an on-chip fabric measured in tens of TB/s, and the Cerebras WSE-3 keeps about 44 GB of SRAM across roughly 900,000 cores at a reported tens-of-petabytes-per-second aggregate bandwidth. That is one to two orders of magnitude more bandwidth per byte than HBM, which crushes decode latency for a model small enough to fit. The catch is capacity: a single 70B model does not fit on one LPU or one wafer, so these systems span a model across many chips and pay the interconnect cost instead. They are compelling for small-to-medium models, low-latency decoding and speculative drafting; they are not a general replacement for an HBM GPU running a 400B model.
TPU and Trainium
The non-NVIDIA column
Google's TPU v5e is a small, cheap serving chip: 16 GB of HBM at 0.8 TB/s, 197 dense BF16 TFLOP/s, and a 400 GB/s bidirectional ICI link — bandwidth-starved per chip, so it is deployed in slices where the ICI fabric carries the sharding. TPU v5p is the opposite: 95 GB at 2.765 TB/s, 459 TFLOP/s BF16, and 1,200 GB/s of ICI. AWS Trainium2 sits in between per chip — 96 GB HBM3 at about 2.9 TB/s and roughly 650 dense BF16 TFLOP/s, with a 16-chip trn2.48xlarge presenting 20.8 FP8 PFLOPS and 1.5 TB of HBM. Treat the interconnect names as the important detail: ICI and NeuronLink are not NVLink, and whether your tensor-parallel group fits inside one ICI domain is the single biggest performance decision on these stacks.
Bytes of weights that fit in HBM, at the selected precision.
Energy per token
Joules, the number the invoice is written in
Energy per token is the accelerator's power draw multiplied by the step time, divided by the number of tokens the step produced. At batch 1 a decode step is bandwidth-bound and its time is fixed, so energy per token is at its worst: you pay for a full weight sweep to produce a single token. As the batch grows the step time is nearly flat until compute takes over, so the same joules now produce many tokens and energy per token collapses. This is the entire economic case for continuous batching in one curve, and it is why a serving node running at batch 1 can be twenty times less energy-efficient than the same node under load.
Animate with play, or drag the batch slider.
Cheat sheet
The numbers that decide decode
| Quantity | Why it matters | Reference points |
|---|---|---|
| HBM bandwidth | Sets decode step time at low batch; tokens/s ≈ BW ÷ model bytes | H100 3.35 TB/s · B200 8 TB/s · NVL72 576 TB/s aggregate |
| HBM capacity | Weights + KV + activations must fit; else shard or quantize | H100 80 GB · H200 141 GB · B200 192 GB · NVL72 13.4 TB |
| NVLink bandwidth | Hides the all-reduce when fast; tax explodes when slow | H100 900 GB/s · B200 1.8 TB/s · NVL72 130 TB/s |
| Dense FLOPs | Only binds once batch pushes the step compute-bound | H100 989 BF16 · B200 2,250 BF16 |
| Fabric between nodes | Carries KV transfer and expert all-to-all | IB NDR 50 GB/s per port · RoCEv2 · NVLink domain 130 TB/s |
| Energy per token | TDP × step time ÷ batch; falls as batch rises | Rack ~120 kW · bandwidth-bound steps waste joules at batch 1 |
Further reading
References
- NVIDIA, H100 Tensor Core GPU datasheet and B200 / GB200 NVL72 architecture pages.
- Pope et al., "Efficiently Scaling Transformer Inference" (MLSys 2023) — the memory-bandwidth model behind decode step time.
- Google Cloud, TPU v5e and v5p documentation; AWS, Trainium2 (trn2.48xlarge) product page.
- Cerebras, "WSE-3" architecture overview; Groq, LPU architecture documentation.