Benchmarks
How the frontier has moved on the evals that matter — general knowledge, graduate-level reasoning, real-world coding, and abstract generalization — release over release.
Figures below are rounded and compiled from public model cards, technical reports, and independent leaderboards (LMSYS Chatbot Arena, Epoch AI, Artificial Analysis, METR). Labs use different evaluation harnesses and prompting setups, so treat exact decimal points as indicative, not authoritative — the trend line is the point.
MMLU — general knowledge
57 subjects spanning STEM, humanities, law and medicine, tested zero- or few-shot. Estimated human-expert baseline is ~89.8%. By mid-2024, frontier models had clustered near that ceiling and labs shifted attention to harder evals below.
| Model | Released | MMLU |
|---|---|---|
| GPT-3 | 2020 | ~44% |
| GPT-4 | Mar 2023 | ~86% |
| Claude 2 | Jul 2023 | ~79% |
| Gemini Ultra | Dec 2023 | ~90% |
| Claude 3 Opus | Mar 2024 | ~87% |
| Llama 3.1 405B | Jul 2024 | ~88% |
| GPT-4o | May 2024 | ~89% |
| Claude 3.5 Sonnet | Jun 2024 | ~89% |
GPQA Diamond — graduate-level reasoning
PhD-level science questions in biology, physics and chemistry, written so domain experts can't answer them by searching. A cleaner test of reasoning than of recall.
| Model | Released | GPQA Diamond |
|---|---|---|
| GPT-4 | Mar 2023 | ~35% |
| Claude 3.5 Sonnet | Jun 2024 | ~59% |
| OpenAI o1 | Sep 2024 | ~78% |
| Claude 3.7 Sonnet | Feb 2025 | ~78% |
| Claude 4 (Opus) | May 2025 | ~83% |
| GPT-5 | Aug 2025 | ~85%+ |
SWE-bench Verified — real-world coding
Resolving genuine GitHub issues end to end: read the repo, write the patch, pass the tests. The closest proxy on this page for "can it do a developer's job," and the benchmark the agentic-coding race has organized around.
| Model | Released | SWE-bench Verified |
|---|---|---|
| Claude 3.5 Sonnet (computer use) | Oct 2024 | ~49% |
| Claude 3.7 Sonnet | Feb 2025 | ~63% |
| Claude 4 (Opus / Sonnet) | May 2025 | ~72–73% |
| GPT-5 | Aug 2025 | ~74% |
| Frontier models per 2026 aggregators | 2026 | ~80%+ |
ARC-AGI — abstract generalization
Visual puzzles solvable by most humans in minutes but historically resistant to pattern-matching by LLMs, since each puzzle is novel rather than drawn from training-adjacent data. Designed specifically to be hard to game.
| Model | Released | ARC-AGI |
|---|---|---|
| GPT-4 | 2023 | ~5% |
| Claude 3.5 Sonnet | 2024 | ~14% |
| OpenAI o3 (preview, high-compute) | Dec 2024 | ~76–88%* |
| Frontier models through 2025 | 2025 | closing on that range, prompting a harder ARC-AGI-2 |
MMMU — multimodal reasoning
College-level questions across six disciplines that require reading a diagram, chart or image alongside the text. Scores are approximate and move with prompting and harness; image-heavy benchmarks are especially sensitive to resolution and token budget, and several items have been found to be answerable from the text alone.
| Model | Released | MMMU |
|---|---|---|
| GPT-4V | 2023 | ~56% |
| Gemini Ultra | Dec 2023 | ~59% |
| Claude 3.5 Sonnet | Jun 2024 | ~68% |
| GPT-4o | 2024 | ~69% |
| OpenAI o1 | Dec 2024 | ~78% |
| Frontier models per 2026 aggregators | 2026 | ~85% |
DocVQA — document understanding
Questions over scanned documents and forms, scored as average normalised Levenshtein similarity (ANLS). The benchmark is close to saturation for frontier multimodal models, so treat any remaining gap as noise and treat high scores as a floor rather than a signal; OCR-heavy training data is now common enough that contamination is a live concern.
| Model | Released | DocVQA |
|---|---|---|
| GPT-4V | 2023 | ~88% ANLS |
| Gemini 1.5 Pro | 2024 | ~93% ANLS |
| Claude 3.5 Sonnet | 2024 | ~95% ANLS |
| Frontier models through 2026 | 2026 | ~96%+ ANLS |
MathVista — visual mathematical reasoning
Mathematical problems that require reading a figure — geometry diagrams, plots, tables and puzzles. Harder to saturate than document QA because the visual structure carries the reasoning, not just the text.
| Model | Released | MathVista |
|---|---|---|
| GPT-4V | 2023 | ~49% |
| Gemini Ultra | 2024 | ~53% |
| GPT-4o | 2024 | ~63% |
| OpenAI o1 | Dec 2024 | ~73% |
| Frontier models through 2026 | 2026 | ~82% |
Text-to-image Elo — human preference
Pairwise human votes on generated images, aggregated into an Elo rating by an arena. It is the closest thing the field has to a human anchor, and it is expensive to maintain — ratings depend on the prompt distribution, the voter population and the comparison budget, and they say nothing about whether a particular prompt was followed. Numbers here are illustrative of the ordering, not a leaderboard snapshot.
| Model | Released | Text-to-image Elo |
|---|---|---|
| Stable Diffusion XL | 2023 | ~1050 Elo |
| Stable Diffusion 3 Medium | 2024 | ~1075 Elo |
| FLUX.1 dev | 2024 | ~1120 Elo |
| Imagen 3 | 2024 | ~1140 Elo |
| GPT-4o image generation | 2025 | ~1150 Elo |
| Leading arena models | 2026 | ~1170 Elo |
*o3's headline number used a high-compute configuration far more expensive per task than typical inference — a reminder that "SOTA" often carries an asterisk about cost.
2025–2026 figures are drawn from AI-news aggregators and leaderboard snapshots rather than primary technical reports, since they fall after most models' training cutoffs — treat them as directionally correct rather than precise.