Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

What the metrics measure

Three numbers, three different questions

FID embeds a set of real images and a set of generated images with an Inception network, fits a Gaussian to each cloud of feature vectors, and reports the Fréchet distance between the two Gaussians, $\mathrm{FID} = \lVert \mu_r - \mu_g \rVert^2 + \mathrm{Tr}\!\left(\Sigma_r + \Sigma_g - 2(\Sigma_r \Sigma_g)^{1/2}\right)$. FVD does the same with a video feature extractor. Both measure the marginal distribution of outputs: how much the cloud of generations looks like the cloud of photographs. Neither is shown the prompt. A model that ignores conditioning entirely and emits a diverse pile of plausible photographs can post a better FID than a model that follows every instruction, because instruction-following changes the conditional distribution while FID only looks at the pooled one. That is the first blind spot, and it is not a bug in the implementation; it is what the metric is.

The second blind spot is that a single scalar is diagnostically empty. A model whose faces melt, one that cannot count, and one that ignores half the prompt can share a score, and the number gives no clue which failure you have. Add the estimator's own noise: FID is biased upward at small sample sizes, it moves with the feature extractor and its resizing and quantisation, and it is not comparable across papers that preprocessed differently. FVD inherits all of it and adds a feature extractor that mixes spatial and temporal quality, so a temporal flicker and a blurry frame land in the same scalar.

CLIPScore asks a different question: it is the cosine $\cos(\phi_\mathrm{img}(x), \phi_\mathrm{text}(c))$ between the image embedding and the text embedding, so it measures alignment, not realism. A cluttered image that contains every noun in the prompt scores highly while humans find it ugly; a beautiful, minimal, on-brief photograph can score lower because it does not show the literal objects. The point cloud below makes the divergence concrete: CLIP is computed for a set of samples, a human-preference proxy is drawn for the same samples, and the two can rank them in opposite orders.

CLIPScore against a human-preference proxy, three prompt families, drawn with GenMedia.plot. The line is the least-squares fit over all samples.

Each point is one generated sample drawn from a seeded stream. The orange family is prompt-literal but cluttered: CLIP likes it, the preference proxy does not. Pooling the three families gives a weak, misleading correlation — here r is about 0.3 — even though both quantities are meaningful inside each family.

💡 By the end of this part you'll be able to say what FID, FVD and CLIPScore each measure and what each cannot see, explain how preference models and arena Elo are built and how they are biased, describe how memorisation is measured and what attribution can and cannot prove, place SynthID, C2PA and the EU marking obligations in relation to one another, and do the serving arithmetic for steps, precision and batch size.
2

Preference models and arenas

Asking a model what a person would have chosen

A preference model is trained on human choices rather than on captions. Pick-a-Pic collected hundreds of thousands of comparisons on user-written prompts, HPS and its successor trained on the HPD v2 dataset, and the resulting scorers — PickScore among them — predict which of two images a person preferred. They correlate with human judgement noticeably better than CLIPScore does, and they are cheap and differentiable, which is exactly why they are dangerous: a differentiable proxy gets optimised, and a generator tuned against PickScore learns PickScore's quirks rather than human taste. The score also inherits the annotator pool, the prompt distribution and the interface that produced the comparisons. It is a statement about a particular population on a particular set of prompts, not a property of the image.

Human arenas take the other route. Users are shown two anonymous outputs to the same prompt, vote, and the votes are fitted with a Bradley–Terry model whose ratings are usually reported on an Elo scale. GenAI-Arena is the generative-media instance of the pattern. Elo is honest about what it is — a ranking inferred from pairwise votes — and it has all the classic fragilities: it is a relative scale that drifts as the pool of models changes, a small prompt set or a self-selected voter population shifts it, ties and abstentions have to be handled by a rule, and a model that wins by being visually striking rather than correct does well unless style is regressed out. The demo below simulates the arithmetic, including a deliberate bias: one model is given a prompt-set advantage it did not earn, and its rating climbs anyway.

Elo fitted from simulated pairwise votes, drawn with Guide.drawBars. Bars start at 1100 so the spread is visible; the labels are the ratings themselves.

Rating sketchTypical valueWhat it is a statement about
FID (image sets)single digits to ~30, lower betterthe pooled distribution of outputs, conditioning ignored
CLIPScore~0.25 to ~0.35 cosinetext–image alignment only, quality ignored
PickScore / HPS v2~0.20 to ~0.24 probabilitythe choices of a specific annotator population
Arena Elo~1100 to ~1300, mean pinnedpairwise votes over a prompt set and a model pool
⚠ An Elo lead is not a quality guarantee. Ratings are relative, they are computed on the prompt set the arena happened to collect, they move when a strong model joins the pool, and any differentiable scorer will be gamed once it is used as a training signal. Report the prompt set, the vote count and the confidence interval, or the number means very little.
3

Memorisation and attribution

When the model outputs a copy

Diffusion models do reproduce their training data. Carlini and coauthors extracted on the order of a hundred near-verbatim training images from Stable Diffusion by prompting for them, and Somepalli and coauthors found that replication concentrates on images that appear many times in the training set and are atypical for their caption. The mechanism is duplication in the data: a picture seen hundreds of times is a low-loss point for the denoiser, and the model can converge to reconstructing it almost exactly. De-duplicating the corpus and conditioning on richer captions both reduce it, which is why data curation is an evaluation topic rather than only a training one.

The measurement is a nearest-neighbour search. Embed the generation, embed the training set, and record the maximum cosine similarity $\max_{z \in \mathcal{D}} \cos(f(x), f(z))$. For an ordinary prompt the nearest neighbour sits well below one; for a memorised prompt the distribution grows a second mode close to one. The demo simulates a hundred prompts against a small training set and lets you raise the effective duplication, which is what a crawled corpus with near-duplicates looks like. Anything above roughly 0.9 in a feature space like this is the region where a lawyer and an auditor both want to look.

Attribution tries to answer the next question: which training items, if any, made this output. Influence functions attempt it by estimating how the loss on an output would change if a training item were removed; retrieval methods simply report the nearest neighbours with their captions; membership-inference methods ask whether a given image was in the training set at all. All of them are noisy and all of them are contested. Deduplication, licensed or opt-out corpora, and provenance recorded at creation time are the practical responses, because a metric that fires after the fact cannot un-train a model.

Distribution of the maximum training-set cosine similarity over 96 prompts, from GenMedia.attn.cosine. Bins past the dashed line are near-duplicates.

The spike near one is the memorised mode. Raising duplication grows it; de-duplication removes it without changing the ordinary prompts.

⚠ A low similarity proves nothing either. Feature-space similarity is not the legal test — substantial similarity of protectable expression is — and a high score is not by itself proof of copying. Treat the number as a triage signal: it tells you which generations deserve a human and a lawyer, not what a court would say.
4

Provenance and regulation

Watermarks live in pixels, manifests live in metadata

SynthID is the pixel-side answer. Rather than appending a tag after generation, it biases the sampling process so that the residual the sampler leaves behind carries a key-dependent, spatially distributed pattern. Detection correlates the image against the key, which is why the signal survives operations that a tag would not: a crop keeps the part of the pattern inside the crop, resizing keeps its frequency structure, and lossy compression attenuates but does not erase it. Formally the detector reports $z = \sqrt{n}\,\mathrm{corr}(v, k)$ for the retained pixels $v$ and the key $k$, so the statistic falls only with the square root of the retained pixel count. The demo below embeds a tiled key into a small procedural image and lets you crop it, then runs the correlation detector on two key layouts — a distributed one and a single corner tag. The distributed signal keeps verifying after a crop; the corner tag dies as soon as its corner is outside the crop.

C2PA is the metadata-side answer, and it is a different trade. Content Credentials are a signed manifest attached to the file describing who created it, with which tool, and what edits were applied. The cryptography is strong: a manifest cannot be forged without the signer's key. It is also fragile in exactly the wrong place, because most social platforms strip metadata on upload, so a screenshot usually arrives with no manifest at all and the absence of a credential says nothing about the content. That is why the two techniques are complements rather than substitutes: a manifest is authoritative while it is attached, and a watermark is weaker but travels with the pixels.

A watermarked procedural image (GenMedia.raster). The dashed square is the retained crop; the small box in the corner is the localised tag.

Detector z-score against the retained fraction, for the distributed key and for the corner tag. Dashed line: the detection threshold. The tag carries far fewer carrier pixels at the same per-pixel amplitude, so its curve sits lower and needs more amplitude to reach the line at all.

Regulation is the third leg, and it is the one with a date on it. Article 50 of the EU AI Act sets transparency duties for providers and deployers of generative systems: providers must mark synthetic outputs in a machine-readable format, and deployers must label deepfakes. Those obligations bind from 2 August 2026. A machine-readable mark can be a watermark or a manifest, which is precisely why the two technical families matter beyond research: one of them has to survive contact with a platform that re-encodes every upload.

⚠ The absence of a watermark is not evidence of authenticity. A generated image can be washed through an edit that destroys the signal, re-photographed, or produced by a model that never embedded one. Provenance is positive evidence only: a verified mark says something; no mark says nothing.
5

Serving economics

What an image actually costs on a GPU

An image is generated by a fixed number of forward passes of a large transformer, and each pass is compute-bound over thousands of latent tokens. The cost of one image is therefore roughly steps times FLOPs-per-step times the price of a GPU-second, divided by how well the hardware is actually used: $\text{cost} \propto s\,F / (\epsilon \cdot \text{peak})$ for $s$ steps, $F$ FLOPs per step and achieved efficiency $\epsilon$. The demo below prices a twelve-billion-parameter backbone at a four-thousand-token latent grid across three configurations, and the shape of the curves is the lesson: at batch one the GPU is under-occupied and per-image cost is high; as the batch grows the fixed weight read is amortised, efficiency rises toward its ceiling, and the curves flatten. Past that knee, more batch buys throughput headroom and latency smoothing rather than a cheaper image, because the arithmetic itself is now the limit.

The levers are the ones the rest of the guide has already built. Fewer steps is the largest: the distillation methods of the few-step part collapse fifty steps into one or four. Step caching is the cheapest: adjacent sampler steps produce similar features, so a cache that reuses a block's activations on the steps where the change is small can skip a substantial fraction of the work at a small quality cost, in the style of DeepCache and TeaCache. Quantization buys both bandwidth and FLOPs — FP8 weights and activations roughly double the arithmetic rate at a quality cost that has to be measured on images, not assumed from language-model results, because activation outliers behave differently in a vision backbone. Batching raises utilisation and absorbs arrival variance.

Cost per thousand images against batch size, from the arithmetic in GenMedia.cost. Three configurations, one vline at the current batch.

Latency

The request-side view

A user waiting on one image cares about time to first pixel, not cost per thousand. Those two numbers trade against each other through batch size: large batches are cheap per image and slow per request. The serving guide's framing of latency budgets, queueing and goodput applies unchanged — LLM Serving, Interactively is the reference for the SLO arithmetic, the percentile discipline and the capacity planning that turns a per-image cost into a fleet size.

Levers

Where the seconds actually go

Steps dominate, so distillation and caching come first. Precision comes second: FP8 or FP4 weights and activations change the arithmetic rate, not the algorithm. Batch size comes third, and it is a utilisation lever with a knee rather than a slope. The few-step part covers the step side and the samplers part covers where the steps come from in the first place.

Further reading

The evaluation literature is a sequence of proxies and their refutations. These are the primary sources for the metrics, the preference datasets and the arenas, the extraction results that made memorisation a first-class concern, and the provenance machinery that regulation now leans on.

Heusel and coauthors introduced FID; Unterthiner and coauthors carried the same construction to video as FVD; Hessel and coauthors proposed the reference-free CLIPScore. Kirstain and coauthors built Pick-a-Pic and PickScore; Wu and coauthors built HPS v2. Carlini and coauthors extracted training images from Stable Diffusion; Somepalli and coauthors characterised replication. On the provenance side, the SynthID papers and the C2PA specification define the two families of marking, and Article 50 of the EU AI Act supplies the deadline.

Cheat sheet

TermMeaning here
FIDFréchet distance between Gaussians of real and generated Inception features; marginal realism only, no conditioning, noisy at small n
FVDThe same construction on video features; spatial and temporal quality collapse into one number
CLIPScoreCosine between CLIP image and text embeddings; alignment, blind to realism, anatomy and aesthetics
PickScore / HPSPreference models trained on human choices; inherit the annotator pool and the prompt set, and are gamed once optimised against
Arena EloBradley–Terry ratings from pairwise human votes; relative, prompt-set dependent, style-sensitive
Nearest-neighbour similarityMaximum cosine to the training set; the memorisation signal that spikes above roughly 0.9
AttributionInfluence functions, retrieval and membership inference; noisy evidence, not a legal conclusion
SynthIDKey-dependent residual embedded during sampling; detected by correlation, survives crops and compression
C2PASigned provenance manifest in metadata; authoritative while attached, stripped by most platforms
Article 50Machine-readable marking of synthetic output, binding from 2 August 2026
Step cachingReusing backbone features across adjacent sampler steps; the cheapest server-side speedup
Serving leversFewer steps, then precision, then batch size; the first is a training decision, the last has a knee
7

Check your understanding

0/4 answered