Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

HBM bandwidth is the headline number

Why spec sheets lead with TB/s, not TFLOP/s

A prefill pass multiplies a whole prompt against the weights at once, so it saturates the tensor cores and is compute-bound. A decode step multiplies one token against the same weights, so it is memory-bandwidth-bound: the arithmetic intensity is roughly one to two FLOPs per byte read, far below the ridge point of any modern accelerator. The practical decode step time is therefore set by how fast HBM can stream the weights in, not by how many FLOPs the chip can issue.

$$t_{\text{step}} \approx \max\!\left(\frac{\text{weight bytes} + \text{KV bytes}}{\text{HBM BW}},\ \frac{2\,N_{\text{active}}}{\text{peak FLOP/s}}\right)$$

The first term barely changes with batch size — the weights are read once per step regardless of how many sequences share that step — which is exactly why batching lifts throughput without a matching latency cost. The second term grows linearly with batch, and where the two cross is the point at which a decode step flips from memory-bound to compute-bound. H100 SXM sits at 3.35 TB/s of HBM and 989 dense BF16 TFLOP/s; B200 doubles the bandwidth to 8 TB/s and the dense throughput to 2,250 BF16 TFLOP/s. The bandwidth doubling is the part that matters for decode.

💡 The rule of thumb: tokens per second per sequence at batch 1 is roughly HBM bandwidth divided by model bytes. A 70B model in BF16 (≈141 GB) on an H100 (3.35 TB/s) tops out near 24 tokens/s; the same model on a B200 (8 TB/s) is closer to 57. Nothing else on the spec sheet moves that number as much.
2

The accelerator zoo

Pick up to four; switch the metric

The comparison below uses dense (non-sparse) throughput where vendors quote sparsity, because serving without 2:4 sparsity runs at the dense rate. Switch the metric and the ranking rearranges violently: NVIDIA leads on HBM bandwidth and NVLink, AMD leads on HBM capacity per dollar-tier, and the TPUs and Trainium trade peak FLOPs for cost and integration inside their own clouds. The decode estimate reads the model's active parameters once per token on each chip, so a 671B-parameter MoE with 37B active looks cheap per token — you pay for the active slice, not the whole file.

Dense figures. Bars are normalized to the largest selected value.

Select up to four accelerators

3

Capacity, and what it forces

Weights + KV + activations versus one HBM budget

Bandwidth decides how fast a step runs; capacity decides whether it runs at all. A single GPU must hold the sharded weights, the KV cache for every resident sequence, and a working set of activations. The weights scale with parameter count and precision, the KV cache scales with context length, concurrency and the attention scheme (Part 4 sizes it properly), and activations scale with batch and hidden width. The moment the stack crosses the HBM capacity the engine has to shard, quantize, or cut context — there is no third option.

The horizontal line is the accelerator's HBM capacity.

5

InfiniBand, RoCE and Ethernet

The fabric between nodes

Inside a node, NVLink carries tensor-parallel collectives. Between nodes, the options are InfiniBand (NDR at 400 Gb/s ≈ 50 GB/s per port, XDR at 800 Gb/s), RoCEv2 (RDMA over Converged Ethernet, which is what most cloud fleets actually run), and plain Ethernet. Their job in serving is different from training: they carry KV transfers between disaggregated prefill and decode pools, expert-parallel all-to-all traffic for MoE, and cache-fetch traffic from a remote KV tier. Two things matter. First, a single KV transfer is a latency-sensitive burst, so RDMA (IB or RoCE) beats TCP/IP by a wide margin. Second, expert all-to-all is bandwidth-bound and scales with the rack's aggregate fabric, which is why NVLink-based racks exist at all.

Relative time to move 1 GB of KV between two nodes.

6

GB200 NVL72 as one big GPU

A 130 TB/s single NVLink domain

GB200 NVL72 wires 72 Blackwell GPUs into one NVLink domain: 1.8 TB/s per GPU, roughly 130 TB/s aggregate, across a liquid-cooled rack drawing on the order of 120 kW. The aggregate is what lets a 671B-parameter model run tensor-parallel and expert-parallel without ever touching Ethernet — the model is effectively one compute unit with 13.4 TB of HBM3e at 576 TB/s aggregate. For serving, that transforms the sharding question: instead of fitting a model onto one 80 GB GPU, you fit it onto a rack and treat the rack as the allocation unit. The cost is that you cannot scale to zero and you pay for the whole rack's power whether it is busy or not.

Per-GPU versus per-rack HBM and bandwidth.

7

Groq and Cerebras: SRAM architectures

When the weights live on-die

The memory wall exists because HBM is off-chip. A different answer is to put the weights in on-die SRAM and give up on scaling a single model: the Groq LPU keeps on the order of 230 MB of SRAM per chip with an on-chip fabric measured in tens of TB/s, and the Cerebras WSE-3 keeps about 44 GB of SRAM across roughly 900,000 cores at a reported tens-of-petabytes-per-second aggregate bandwidth. That is one to two orders of magnitude more bandwidth per byte than HBM, which crushes decode latency for a model small enough to fit. The catch is capacity: a single 70B model does not fit on one LPU or one wafer, so these systems span a model across many chips and pay the interconnect cost instead. They are compelling for small-to-medium models, low-latency decoding and speculative drafting; they are not a general replacement for an HBM GPU running a 400B model.

💡 The trade in one line: SRAM buys bandwidth at the cost of capacity, HBM buys capacity at the cost of bandwidth. Serving economics is the art of picking which one your workload is short of.
8

TPU and Trainium

The non-NVIDIA column

Google's TPU v5e is a small, cheap serving chip: 16 GB of HBM at 0.8 TB/s, 197 dense BF16 TFLOP/s, and a 400 GB/s bidirectional ICI link — bandwidth-starved per chip, so it is deployed in slices where the ICI fabric carries the sharding. TPU v5p is the opposite: 95 GB at 2.765 TB/s, 459 TFLOP/s BF16, and 1,200 GB/s of ICI. AWS Trainium2 sits in between per chip — 96 GB HBM3 at about 2.9 TB/s and roughly 650 dense BF16 TFLOP/s, with a 16-chip trn2.48xlarge presenting 20.8 FP8 PFLOPS and 1.5 TB of HBM. Treat the interconnect names as the important detail: ICI and NeuronLink are not NVLink, and whether your tensor-parallel group fits inside one ICI domain is the single biggest performance decision on these stacks.

Bytes of weights that fit in HBM, at the selected precision.

9

Energy per token

Joules, the number the invoice is written in

Energy per token is the accelerator's power draw multiplied by the step time, divided by the number of tokens the step produced. At batch 1 a decode step is bandwidth-bound and its time is fixed, so energy per token is at its worst: you pay for a full weight sweep to produce a single token. As the batch grows the step time is nearly flat until compute takes over, so the same joules now produce many tokens and energy per token collapses. This is the entire economic case for continuous batching in one curve, and it is why a serving node running at batch 1 can be twenty times less energy-efficient than the same node under load.

Animate with play, or drag the batch slider.

📌 Cross-link: the memory accounting here — weights plus optimizer state during training — is exactly the arithmetic worked through in the LLM Training scaling page, where the same parameter is multiplied by 16 bytes instead of 2.
✓

Cheat sheet

The numbers that decide decode

QuantityWhy it mattersReference points
HBM bandwidthSets decode step time at low batch; tokens/s ≈ BW ÷ model bytesH100 3.35 TB/s · B200 8 TB/s · NVL72 576 TB/s aggregate
HBM capacityWeights + KV + activations must fit; else shard or quantizeH100 80 GB · H200 141 GB · B200 192 GB · NVL72 13.4 TB
NVLink bandwidthHides the all-reduce when fast; tax explodes when slowH100 900 GB/s · B200 1.8 TB/s · NVL72 130 TB/s
Dense FLOPsOnly binds once batch pushes the step compute-boundH100 989 BF16 · B200 2,250 BF16
Fabric between nodesCarries KV transfer and expert all-to-allIB NDR 50 GB/s per port · RoCEv2 · NVLink domain 130 TB/s
Energy per tokenTDP × step time ÷ batch; falls as batch risesRack ~120 kW · bandwidth-bound steps waste joules at batch 1
📚

Further reading

References

?

Check your understanding

0/5 answered