THE ARTICLE · 9 MIN
When a company releases an AI model, it usually publishes a table of benchmark scores to show how the model compares with others. Those numbers are useful, but they are easy to over-read. This page explains what the best-known benchmarks test and the specific, documented ways a score can say less than it seems. Scores and policies change quickly; this is a snapshot as of September 2026.
What a benchmark is
A benchmark is a fixed set of tasks with known answers, scored automatically, so that different models can be measured on exactly the same questions.
A classic example is MMLU. Its authors wrote: “The test covers 57 tasks including elementary mathematics, US history, computer science, law, and more.” When it was published, they found that “most recent models have near random-chance accuracy”.
Tests like this stop being useful once the best models score near the top. The authors of a newer benchmark, Humanity’s Last Exam, wrote that “benchmarks are not keeping pace in difficulty”, noting that models now score over 90% on popular benchmarks like MMLU. This is called saturation: when nearly every leading model passes, the test can no longer tell them apart.
The benchmarks you will see most often
| Benchmark | What it tests | A number worth knowing |
|---|---|---|
| GPQA | “448 multiple-choice questions written by domain experts in biology, physics, and chemistry” | Experts with or pursuing PhDs reached 65% (74% after discounting mistakes they spotted afterwards); skilled non-experts reached 34% even with web access. |
| Humanity’s Last Exam (HLE) | Hard academic questions; “HLE consists of 2,500 questions across dozens of subjects, including mathematics, humanities, and the natural sciences.” | Each question “has a known solution that is unambiguous and easily verifiable, but cannot be quickly answered via internet retrieval.” |
| SWE-bench | Real software problems: “Given a codebase along with a description of an issue to be resolved, a language model is tasked with editing the codebase to address the issue.” | 2,294 problems from 12 popular Python repositories; at launch the best model solved 1.96%. |
| SWE-bench Verified | A smaller set of SWE-bench problems checked by people: “a subset of the original test set from SWE-bench, consisting of 500 samples verified to be non-problematic by our human annotators” | Released by OpenAI with the SWE-bench authors in August 2024. See below for why OpenAI later stopped using it. |
| ARC-AGI-3 | Interactive puzzles: “There are no instructions, no rules, and no stated goals.” | At its launch, the organisers reported: “Humans score 100%. Frontier AI scores 0.51%.” |
| METR time horizon | How long a task a model can handle: “the time humans typically take to complete tasks that AI models can complete with 50% success rate” | METR’s paper reports this “has been doubling approximately every seven months since 2019, though the trend may have accelerated in 2024”. A January 2026 update estimated a doubling time of about 131 days for models since 2023. |
Why a score can mislead
1. The test questions leak into training data
AI models are trained on huge amounts of text gathered from the internet, and benchmark questions are often published there too. Researchers call this data contamination: “the unintended overlap between training and test datasets”, which can lead to “an overestimation of the models’ true generalization capabilities.”
The clearest example comes from a benchmark’s own co-creator. In February 2026 OpenAI explained why it had dropped SWE-bench Verified. It said that “all frontier models we tested were able to reproduce the original, human-written bug fix used as the ground-truth reference, known as the gold patch, or verbatim problem statement specifics for certain tasks” — in other words, each model had seen at least some of the problems and their solutions during training. It concluded: “This is why we have stopped reporting SWE-bench Verified scores, and we recommend that other model developers do so too.”
2. The answer key can be wrong
The same OpenAI analysis found a second problem. It “audited a 27.6% subset of the dataset that models often failed to solve and found that at least 59.4% of the audited problems have flawed test cases that reject functionally correct submissions”.
That figure applies to the audited problems, not to the whole benchmark. But it shows how a model can be marked wrong for a correct answer — and why a score near the top of a test may say more about the test than about the model.
3. The way the test is run changes the result
OpenAI put it plainly in July 2026: “Benchmarks rarely measure AI models in isolation.” A model runs inside a harness — the software that feeds it the task, keeps track of its progress and collects its answer — and often at a chosen level of reasoning effort.
On 3 September 2026, ARC Prize published its results for OpenAI’s GPT-6 Astra on the Semi-Private test set of ARC-AGI-3. With its Standard harness the model scored 62.7% at maximum effort and 54.8% at high effort. With a Provider Adapter harness — which “preserves opaque reasoning state between requests and uses compaction for longer conversations, allowing the model to reuse prior work” — it scored 99.9% at high effort and 98.6% at maximum effort. The organisers’ summary put two of those runs side by side: “62.7% for $26K” and “99.9% for $19K”.
Same model, same benchmark, very different scores. A benchmark number without its setup is only half a fact.
4. The tasks may not be the tasks you care about
METR, which publishes the time-horizon measurements, lists the limits of its own benchmark. It says the time horizon is “a measure of the difficulty of a task, rather than the time an AI spends to complete the task.” It says: “Our task distribution is primarily composed of software engineering, machine learning, or cybersecurity tasks.” It warns: “In other words, our tasks are much “cleaner” than real economically valuable labor.” And its chart notes: “Measurements above 16 hrs are unreliable with our current task suite”.
A high score on a clean, well-specified software task does not show how a model will do on a messy task in a different field.
5. Who gets to test before release
Some leaderboards rank models by public votes rather than fixed questions. The best known is Arena (formerly Chatbot Arena, then LMArena), where people compare answers from two anonymous models and pick the better one. A 2025 paper, “The Leaderboard Illusion”, argued that its rules gave some developers an advantage. Its operators, then called LMArena, published a response disputing several of its claims.
Disputed The two sides agree that developers tested private versions of models before release. They disagree about whether that practice was undisclosed or favoured some developers, and about how much it distorted the rankings.
- The paper reported that “undisclosed private testing practices benefit a handful of providers who are able to test multiple variants before public release and retract scores if desired”. It estimated that two large providers received 19.2% and 20.4% of all the arena’s data, while “a combined 83 open-weight models have only received an estimated 29.7% of the total data”, and reported “relative performance gains of up to 112% on ArenaHard, a test set from the arena distribution”, from training on arena data.
- LMArena replied that its testing policy had been public since March 2024 and that “any model provider can submit as many public and private variants as they would like, as long as we have capacity for it”, while also saying that a future policy release would “explicitly state that model providers are all allowed to test multiple variants of their models pre-release”; that ArenaHard is “a static benchmark with 500 data points that uses an LLM judge, and no human labels”, and so not representative of the arena itself; and that its analysis put the effect of pre-release testing at “around +11 Elo after 50 tests and 3000 votes”, falling towards zero as new votes arrive.
Arena’s current policy, last updated in September 2026, says that if a model was tested before release, its score is marked “as preliminary until enough fresh votes have been collected after the model’s public release”.
The evidence does not settle these questions. The safe reading is that a public-vote ranking is one signal among several, not a final verdict.
6. A single “intelligence” number is a choice of ingredients
Some trackers combine many benchmarks into one index. Artificial Analysis says its “Intelligence Index v4.3 incorporates 10 evaluations”, and adds: “Like all evaluation metrics, it has limitations and may not apply directly to every use case.” Its methodology page now lists GPQA Diamond and MMLU-Pro under “Legacy Evaluations”.
Change the ingredients and the ranking can change. A combined score is useful for a quick comparison, but it reflects the tracker’s choice of which tests count.
The old rule behind all of this
The social anthropologist Marilyn Strathern wrote in 1997: “When a measure becomes a target, it ceases to be a good measure.” She drew on the work of an educationalist, Hoskin: “Hoskin describes this as ‘Goodhart’s law’, after the latter’s observation on instruments for monetary control”. Strathern was writing about exam grades in British universities, not AI — but the pattern is the same. Once a benchmark becomes the number everyone competes on, it becomes easier to improve the number than the ability it was meant to measure.
How to read a benchmark claim in two minutes
- Which benchmark, and which version? SWE-bench and SWE-bench Verified are different tests; so are ARC-AGI-2 and ARC-AGI-3.
- Who ran it? The model’s developer, the benchmark’s organisers, or an independent tracker. Results gathered by others are worth comparing: Epoch AI’s hub, for example, “includes results from benchmarks administered internally by Epoch AI as well as data collected from external sources.”
- With what setup? The harness and the reasoning effort can change the result.
- At what cost? Two scores reached at very different cost are not the same achievement.
- Is the test saturated? If every leading model scores near the top, small differences mean little.
- Could it be contaminated? OpenAI noted that SWE-bench Verified was “open-source and broadly used and discussed, which makes avoiding contamination difficult”. Widely published tests carry that risk.
- Are the tasks like yours? A coding benchmark tells you about coding, not about writing or research.
Our reading: benchmarks are the best shared measuring stick the field has, and they do track real progress. But a single score is a result under particular conditions, not a general measure of how capable a model is. Read the conditions before the number.
Sources
- Hendrycks et al., “Measuring Massive Multitask Language Understanding” (MMLU), arXiv:2009.03300.
- Rein et al., “GPQA: A Graduate-Level Google-Proof Q&A Benchmark”, arXiv:2311.12022.
- Phan et al., “Humanity’s Last Exam”, arXiv:2501.14249.
- Jimenez et al., “SWE-bench: Can Language Models Resolve Real-World GitHub Issues?”, arXiv:2310.06770.
- OpenAI, “Introducing SWE-bench Verified” (13 August 2024); “Why SWE-bench Verified no longer measures frontier coding capabilities” (23 February 2026).
- ARC Prize, “Announcing ARC-AGI-3” (25 March 2026); “OpenAI’s GPT-6 Astra on ARC-AGI-3” (3 September 2026).
- Kwa et al., “Measuring AI Ability to Complete Long Software Tasks”, arXiv:2503.14499; METR, “Time horizons” page; METR, “Time Horizon 1.1” (January 2026).
- “A Survey on Data Contamination for Large Language Models”, arXiv:2502.14425.
- Singh et al., “The Leaderboard Illusion”, arXiv:2504.20879; LMArena, response to “The Leaderboard Illusion”, arena.ai.
- OpenAI, “How enabling two settings tripled our scores on the ARC-AGI-3 benchmark” (29 July 2026).
- Arena, Leaderboard Policy (last updated 1 September 2026), arena.ai.
- Artificial Analysis, Intelligence Index methodology; Epoch AI, Benchmarking Hub.
- Strathern, M., “‘Improving ratings’: audit in the British University system”, European Review 5(3), 1997, p. 308.
Checked September 2026.
Related: What is a frontier AI model? · Open-weight is not open-source · Numbers that mislead · How to read a study without a science degree
- artificial intelligence
- ai models
- benchmarks
- evaluation
- explainer
