Evaluating and shipping
A generative model produces a picture, and a picture is not a number. Everything in this guide so far was measurable: a loss, a schedule, a token count, a step size. The moment the output is an image or a clip, the honest measurements get harder, and the field's substitutes for them — FID, FVD, CLIPScore, preference models, arena Elo — are each a proxy with a specific blind spot. This part is about those blind spots, about the images a model reproduces rather than generates, about the provenance machinery that is supposed to say where a picture came from, and about the serving arithmetic that decides whether any of it is affordable.
What the metrics measure
Three numbers, three different questions
FID embeds a set of real images and a set of generated images with an Inception network, fits a Gaussian to each cloud of feature vectors, and reports the Fréchet distance between the two Gaussians, $\mathrm{FID} = \lVert \mu_r - \mu_g \rVert^2 + \mathrm{Tr}\!\left(\Sigma_r + \Sigma_g - 2(\Sigma_r \Sigma_g)^{1/2}\right)$. FVD does the same with a video feature extractor. Both measure the marginal distribution of outputs: how much the cloud of generations looks like the cloud of photographs. Neither is shown the prompt. A model that ignores conditioning entirely and emits a diverse pile of plausible photographs can post a better FID than a model that follows every instruction, because instruction-following changes the conditional distribution while FID only looks at the pooled one. That is the first blind spot, and it is not a bug in the implementation; it is what the metric is.
The second blind spot is that a single scalar is diagnostically empty. A model whose faces melt, one that cannot count, and one that ignores half the prompt can share a score, and the number gives no clue which failure you have. Add the estimator's own noise: FID is biased upward at small sample sizes, it moves with the feature extractor and its resizing and quantisation, and it is not comparable across papers that preprocessed differently. FVD inherits all of it and adds a feature extractor that mixes spatial and temporal quality, so a temporal flicker and a blurry frame land in the same scalar.
CLIPScore asks a different question: it is the cosine $\cos(\phi_\mathrm{img}(x), \phi_\mathrm{text}(c))$ between the image embedding and the text embedding, so it measures alignment, not realism. A cluttered image that contains every noun in the prompt scores highly while humans find it ugly; a beautiful, minimal, on-brief photograph can score lower because it does not show the literal objects. The point cloud below makes the divergence concrete: CLIP is computed for a set of samples, a human-preference proxy is drawn for the same samples, and the two can rank them in opposite orders.
CLIPScore against a human-preference proxy, three prompt families, drawn with GenMedia.plot. The line is the least-squares fit over all samples.
Each point is one generated sample drawn from a seeded stream. The orange family is prompt-literal but cluttered: CLIP likes it, the preference proxy does not. Pooling the three families gives a weak, misleading correlation — here r is about 0.3 — even though both quantities are meaningful inside each family.
Preference models and arenas
Asking a model what a person would have chosen
A preference model is trained on human choices rather than on captions. Pick-a-Pic collected hundreds of thousands of comparisons on user-written prompts, HPS and its successor trained on the HPD v2 dataset, and the resulting scorers — PickScore among them — predict which of two images a person preferred. They correlate with human judgement noticeably better than CLIPScore does, and they are cheap and differentiable, which is exactly why they are dangerous: a differentiable proxy gets optimised, and a generator tuned against PickScore learns PickScore's quirks rather than human taste. The score also inherits the annotator pool, the prompt distribution and the interface that produced the comparisons. It is a statement about a particular population on a particular set of prompts, not a property of the image.
Human arenas take the other route. Users are shown two anonymous outputs to the same prompt, vote, and the votes are fitted with a Bradley–Terry model whose ratings are usually reported on an Elo scale. GenAI-Arena is the generative-media instance of the pattern. Elo is honest about what it is — a ranking inferred from pairwise votes — and it has all the classic fragilities: it is a relative scale that drifts as the pool of models changes, a small prompt set or a self-selected voter population shifts it, ties and abstentions have to be handled by a rule, and a model that wins by being visually striking rather than correct does well unless style is regressed out. The demo below simulates the arithmetic, including a deliberate bias: one model is given a prompt-set advantage it did not earn, and its rating climbs anyway.
Elo fitted from simulated pairwise votes, drawn with Guide.drawBars. Bars start at 1100 so the spread is visible; the labels are the ratings themselves.
| Rating sketch | Typical value | What it is a statement about |
|---|---|---|
| FID (image sets) | single digits to ~30, lower better | the pooled distribution of outputs, conditioning ignored |
| CLIPScore | ~0.25 to ~0.35 cosine | text–image alignment only, quality ignored |
| PickScore / HPS v2 | ~0.20 to ~0.24 probability | the choices of a specific annotator population |
| Arena Elo | ~1100 to ~1300, mean pinned | pairwise votes over a prompt set and a model pool |
Memorisation and attribution
When the model outputs a copy
Diffusion models do reproduce their training data. Carlini and coauthors extracted on the order of a hundred near-verbatim training images from Stable Diffusion by prompting for them, and Somepalli and coauthors found that replication concentrates on images that appear many times in the training set and are atypical for their caption. The mechanism is duplication in the data: a picture seen hundreds of times is a low-loss point for the denoiser, and the model can converge to reconstructing it almost exactly. De-duplicating the corpus and conditioning on richer captions both reduce it, which is why data curation is an evaluation topic rather than only a training one.
The measurement is a nearest-neighbour search. Embed the generation, embed the training set, and record the maximum cosine similarity $\max_{z \in \mathcal{D}} \cos(f(x), f(z))$. For an ordinary prompt the nearest neighbour sits well below one; for a memorised prompt the distribution grows a second mode close to one. The demo simulates a hundred prompts against a small training set and lets you raise the effective duplication, which is what a crawled corpus with near-duplicates looks like. Anything above roughly 0.9 in a feature space like this is the region where a lawyer and an auditor both want to look.
Attribution tries to answer the next question: which training items, if any, made this output. Influence functions attempt it by estimating how the loss on an output would change if a training item were removed; retrieval methods simply report the nearest neighbours with their captions; membership-inference methods ask whether a given image was in the training set at all. All of them are noisy and all of them are contested. Deduplication, licensed or opt-out corpora, and provenance recorded at creation time are the practical responses, because a metric that fires after the fact cannot un-train a model.
Distribution of the maximum training-set cosine similarity over 96 prompts, from GenMedia.attn.cosine. Bins past the dashed line are near-duplicates.
The spike near one is the memorised mode. Raising duplication grows it; de-duplication removes it without changing the ordinary prompts.
Provenance and regulation
Watermarks live in pixels, manifests live in metadata
SynthID is the pixel-side answer. Rather than appending a tag after generation, it biases the sampling process so that the residual the sampler leaves behind carries a key-dependent, spatially distributed pattern. Detection correlates the image against the key, which is why the signal survives operations that a tag would not: a crop keeps the part of the pattern inside the crop, resizing keeps its frequency structure, and lossy compression attenuates but does not erase it. Formally the detector reports $z = \sqrt{n}\,\mathrm{corr}(v, k)$ for the retained pixels $v$ and the key $k$, so the statistic falls only with the square root of the retained pixel count. The demo below embeds a tiled key into a small procedural image and lets you crop it, then runs the correlation detector on two key layouts — a distributed one and a single corner tag. The distributed signal keeps verifying after a crop; the corner tag dies as soon as its corner is outside the crop.
C2PA is the metadata-side answer, and it is a different trade. Content Credentials are a signed manifest attached to the file describing who created it, with which tool, and what edits were applied. The cryptography is strong: a manifest cannot be forged without the signer's key. It is also fragile in exactly the wrong place, because most social platforms strip metadata on upload, so a screenshot usually arrives with no manifest at all and the absence of a credential says nothing about the content. That is why the two techniques are complements rather than substitutes: a manifest is authoritative while it is attached, and a watermark is weaker but travels with the pixels.
A watermarked procedural image (GenMedia.raster). The dashed square is the retained crop; the small box in the corner is the localised tag.
Detector z-score against the retained fraction, for the distributed key and for the corner tag. Dashed line: the detection threshold. The tag carries far fewer carrier pixels at the same per-pixel amplitude, so its curve sits lower and needs more amplitude to reach the line at all.
Regulation is the third leg, and it is the one with a date on it. Article 50 of the EU AI Act sets transparency duties for providers and deployers of generative systems: providers must mark synthetic outputs in a machine-readable format, and deployers must label deepfakes. Those obligations bind from 2 August 2026. A machine-readable mark can be a watermark or a manifest, which is precisely why the two technical families matter beyond research: one of them has to survive contact with a platform that re-encodes every upload.
Serving economics
What an image actually costs on a GPU
An image is generated by a fixed number of forward passes of a large transformer, and each pass is compute-bound over thousands of latent tokens. The cost of one image is therefore roughly steps times FLOPs-per-step times the price of a GPU-second, divided by how well the hardware is actually used: $\text{cost} \propto s\,F / (\epsilon \cdot \text{peak})$ for $s$ steps, $F$ FLOPs per step and achieved efficiency $\epsilon$. The demo below prices a twelve-billion-parameter backbone at a four-thousand-token latent grid across three configurations, and the shape of the curves is the lesson: at batch one the GPU is under-occupied and per-image cost is high; as the batch grows the fixed weight read is amortised, efficiency rises toward its ceiling, and the curves flatten. Past that knee, more batch buys throughput headroom and latency smoothing rather than a cheaper image, because the arithmetic itself is now the limit.
The levers are the ones the rest of the guide has already built. Fewer steps is the largest: the distillation methods of the few-step part collapse fifty steps into one or four. Step caching is the cheapest: adjacent sampler steps produce similar features, so a cache that reuses a block's activations on the steps where the change is small can skip a substantial fraction of the work at a small quality cost, in the style of DeepCache and TeaCache. Quantization buys both bandwidth and FLOPs — FP8 weights and activations roughly double the arithmetic rate at a quality cost that has to be measured on images, not assumed from language-model results, because activation outliers behave differently in a vision backbone. Batching raises utilisation and absorbs arrival variance.
Cost per thousand images against batch size, from the arithmetic in GenMedia.cost. Three configurations, one vline at the current batch.
The request-side view
A user waiting on one image cares about time to first pixel, not cost per thousand. Those two numbers trade against each other through batch size: large batches are cheap per image and slow per request. The serving guide's framing of latency budgets, queueing and goodput applies unchanged — LLM Serving, Interactively is the reference for the SLO arithmetic, the percentile discipline and the capacity planning that turns a per-image cost into a fleet size.
Where the seconds actually go
Steps dominate, so distillation and caching come first. Precision comes second: FP8 or FP4 weights and activations change the arithmetic rate, not the algorithm. Batch size comes third, and it is a utilisation lever with a knee rather than a slope. The few-step part covers the step side and the samplers part covers where the steps come from in the first place.
Further reading
The evaluation literature is a sequence of proxies and their refutations. These are the primary sources for the metrics, the preference datasets and the arenas, the extraction results that made memorisation a first-class concern, and the provenance machinery that regulation now leans on.
Heusel and coauthors introduced FID; Unterthiner and coauthors carried the same construction to video as FVD; Hessel and coauthors proposed the reference-free CLIPScore. Kirstain and coauthors built Pick-a-Pic and PickScore; Wu and coauthors built HPS v2. Carlini and coauthors extracted training images from Stable Diffusion; Somepalli and coauthors characterised replication. On the provenance side, the SynthID papers and the C2PA specification define the two families of marking, and Article 50 of the EU AI Act supplies the deadline.
- Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, Sepp Hochreiter, "GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium", 2017 — the FID metric.
- Thomas Unterthiner, Sjoerd van Steenkiste, Karol Kurach, Raphaël Marinier, Marcin Michalski, Sylvain Gelly, "Towards Accurate Generative Models of Video: A New Metric & Challenges", 2018 — FVD.
- Jack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras, Yejin Choi, "CLIPScore: A Reference-free Evaluation Metric for Image Captioning", 2021 — the alignment metric this part criticises.
- Yuval Kirstain, Adam Polyak, Uriel Singer, Shahbuland Matiana, Joe Penna, Omer Levy, "Pick-a-Pic: An Open Dataset of User Preferences for Text-to-Image Generation", 2023 — the preference dataset behind PickScore.
- Xiaoshi Wu, Yiming Hao, Keqiang Sun, Yixiong Chen, Feng Zhu, Rui Zhao, Hongsheng Li, "Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis", 2023 — HPS v2 and HPD v2.
- Nicholas Carlini, Jamie Hayes, Milad Nasr, Matthew Jagielski, Vikash Sehwag, Florian Tramèr, Borja Balle, Daphne Ippolito, Eric Wallace, "Extracting Training Data from Diffusion Models", 2023 — near-verbatim extraction.
- Gowthami Somepalli, Vasu Singla, Micah Goldblum, Jonas Geiping, Tom Goldstein, "Diffusion Art or Digital Forgery? Investigating Data Replication in Diffusion Models", 2022 — where replication comes from.
- Sander Dieleman, "Guidance: a cheat code for diffusion models" and the SynthID-Image technical report from Google DeepMind, 2023–2024 — the sampling-time watermark this part's demo models.
- Coalition for Content Provenance and Authenticity, "C2PA Technical Specification" and the text of Regulation (EU) 2024/1689, Article 50 — content credentials and the marking obligation that binds from 2 August 2026.
Cheat sheet
| Term | Meaning here |
|---|---|
| FID | Fréchet distance between Gaussians of real and generated Inception features; marginal realism only, no conditioning, noisy at small n |
| FVD | The same construction on video features; spatial and temporal quality collapse into one number |
| CLIPScore | Cosine between CLIP image and text embeddings; alignment, blind to realism, anatomy and aesthetics |
| PickScore / HPS | Preference models trained on human choices; inherit the annotator pool and the prompt set, and are gamed once optimised against |
| Arena Elo | Bradley–Terry ratings from pairwise human votes; relative, prompt-set dependent, style-sensitive |
| Nearest-neighbour similarity | Maximum cosine to the training set; the memorisation signal that spikes above roughly 0.9 |
| Attribution | Influence functions, retrieval and membership inference; noisy evidence, not a legal conclusion |
| SynthID | Key-dependent residual embedded during sampling; detected by correlation, survives crops and compression |
| C2PA | Signed provenance manifest in metadata; authoritative while attached, stripped by most platforms |
| Article 50 | Machine-readable marking of synthetic output, binding from 2 August 2026 |
| Step caching | Reusing backbone features across adjacent sampler steps; the cheapest server-side speedup |
| Serving levers | Fewer steps, then precision, then batch size; the first is a training decision, the last has a knee |