Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

Figures below are rounded and compiled from public model cards, technical reports, and independent leaderboards (LMSYS Chatbot Arena, Epoch AI, Artificial Analysis, METR). Labs use different evaluation harnesses and prompting setups, so treat exact decimal points as indicative, not authoritative — the trend line is the point.

MMLU — general knowledge

57 subjects spanning STEM, humanities, law and medicine, tested zero- or few-shot. Estimated human-expert baseline is ~89.8%. By mid-2024, frontier models had clustered near that ceiling and labs shifted attention to harder evals below.

ModelReleasedMMLU
GPT-32020~44%
GPT-4Mar 2023~86%
Claude 2Jul 2023~79%
Gemini UltraDec 2023~90%
Claude 3 OpusMar 2024~87%
Llama 3.1 405BJul 2024~88%
GPT-4oMay 2024~89%
Claude 3.5 SonnetJun 2024~89%

GPQA Diamond — graduate-level reasoning

PhD-level science questions in biology, physics and chemistry, written so domain experts can't answer them by searching. A cleaner test of reasoning than of recall.

ModelReleasedGPQA Diamond
GPT-4Mar 2023~35%
Claude 3.5 SonnetJun 2024~59%
OpenAI o1Sep 2024~78%
Claude 3.7 SonnetFeb 2025~78%
Claude 4 (Opus)May 2025~83%
GPT-5Aug 2025~85%+

SWE-bench Verified — real-world coding

Resolving genuine GitHub issues end to end: read the repo, write the patch, pass the tests. The closest proxy on this page for "can it do a developer's job," and the benchmark the agentic-coding race has organized around.

ModelReleasedSWE-bench Verified
Claude 3.5 Sonnet (computer use)Oct 2024~49%
Claude 3.7 SonnetFeb 2025~63%
Claude 4 (Opus / Sonnet)May 2025~72–73%
GPT-5Aug 2025~74%
Frontier models per 2026 aggregators2026~80%+

ARC-AGI — abstract generalization

Visual puzzles solvable by most humans in minutes but historically resistant to pattern-matching by LLMs, since each puzzle is novel rather than drawn from training-adjacent data. Designed specifically to be hard to game.

ModelReleasedARC-AGI
GPT-42023~5%
Claude 3.5 Sonnet2024~14%
OpenAI o3 (preview, high-compute)Dec 2024~76–88%*
Frontier models through 20252025closing on that range, prompting a harder ARC-AGI-2

MMMU — multimodal reasoning

College-level questions across six disciplines that require reading a diagram, chart or image alongside the text. Scores are approximate and move with prompting and harness; image-heavy benchmarks are especially sensitive to resolution and token budget, and several items have been found to be answerable from the text alone.

ModelReleasedMMMU
GPT-4V2023~56%
Gemini UltraDec 2023~59%
Claude 3.5 SonnetJun 2024~68%
GPT-4o2024~69%
OpenAI o1Dec 2024~78%
Frontier models per 2026 aggregators2026~85%

DocVQA — document understanding

Questions over scanned documents and forms, scored as average normalised Levenshtein similarity (ANLS). The benchmark is close to saturation for frontier multimodal models, so treat any remaining gap as noise and treat high scores as a floor rather than a signal; OCR-heavy training data is now common enough that contamination is a live concern.

ModelReleasedDocVQA
GPT-4V2023~88% ANLS
Gemini 1.5 Pro2024~93% ANLS
Claude 3.5 Sonnet2024~95% ANLS
Frontier models through 20262026~96%+ ANLS

MathVista — visual mathematical reasoning

Mathematical problems that require reading a figure — geometry diagrams, plots, tables and puzzles. Harder to saturate than document QA because the visual structure carries the reasoning, not just the text.

ModelReleasedMathVista
GPT-4V2023~49%
Gemini Ultra2024~53%
GPT-4o2024~63%
OpenAI o1Dec 2024~73%
Frontier models through 20262026~82%

Text-to-image Elo — human preference

Pairwise human votes on generated images, aggregated into an Elo rating by an arena. It is the closest thing the field has to a human anchor, and it is expensive to maintain — ratings depend on the prompt distribution, the voter population and the comparison budget, and they say nothing about whether a particular prompt was followed. Numbers here are illustrative of the ordering, not a leaderboard snapshot.

ModelReleasedText-to-image Elo
Stable Diffusion XL2023~1050 Elo
Stable Diffusion 3 Medium2024~1075 Elo
FLUX.1 dev2024~1120 Elo
Imagen 32024~1140 Elo
GPT-4o image generation2025~1150 Elo
Leading arena models2026~1170 Elo

*o3's headline number used a high-compute configuration far more expensive per task than typical inference — a reminder that "SOTA" often carries an asterisk about cost.

2025–2026 figures are drawn from AI-news aggregators and leaderboard snapshots rather than primary technical reports, since they fall after most models' training cutoffs — treat them as directionally correct rather than precise.