HACKS VITAE

TOOLS & TECH · SOURCES SHOWN

AI Benchmarks Explained
What the Scores Measure and Why They Can Mislead

PRICES, VERSIONS AND FACTS AS OF SEPTEMBER 2026

Every new AI model arrives with a table of benchmark scores. Here is what the best-known benchmarks test, and the specific ways a score can say less than it seems — from leaked test questions to broken answer keys and test setups that change the result.

5 SECTIONS · HOVER A POINT TO JUMP
Published
September 17, 2026
Updated
September 25, 2026
Facts as of
September 2026
Read
9 min
Sections
5

BACKGROUND · REMBRANDT, HERMAN DOOMER, 1640 · THE MET, OPEN ACCESS

THE SHORT VERSION

  1. A benchmark is a fixed test with known answers, scored automatically, so different AI models can be compared on the same questions.
  2. Scores can mislead in specific, documented ways: the test gets too easy, its questions leak into training data, its answer key is wrong, or the way the test is run changes the result.
  3. OpenAI stopped reporting scores on SWE-bench Verified, a coding benchmark it helped create, after finding flawed tests and signs that models had seen the problems during training.
  4. Before trusting a score, ask: which version of the test, who ran it, with what setup and at what cost, and what kind of tasks it contains.

THE ARTICLE · 9 MIN

When a company releases an AI model, it usually publishes a table of benchmark scores to show how the model compares with others. Those numbers are useful, but they are easy to over-read. This page explains what the best-known benchmarks test and the specific, documented ways a score can say less than it seems. Scores and policies change quickly; this is a snapshot as of September 2026.

What a benchmark is

A benchmark is a fixed set of tasks with known answers, scored automatically, so that different models can be measured on exactly the same questions.

A classic example is MMLU. Its authors wrote: “The test covers 57 tasks including elementary mathematics, US history, computer science, law, and more.” When it was published, they found that “most recent models have near random-chance accuracy”.

Tests like this stop being useful once the best models score near the top. The authors of a newer benchmark, Humanity’s Last Exam, wrote that “benchmarks are not keeping pace in difficulty”, noting that models now score over 90% on popular benchmarks like MMLU. This is called saturation: when nearly every leading model passes, the test can no longer tell them apart.

The benchmarks you will see most often

BenchmarkWhat it testsA number worth knowing
GPQA“448 multiple-choice questions written by domain experts in biology, physics, and chemistry”Experts with or pursuing PhDs reached 65% (74% after discounting mistakes they spotted afterwards); skilled non-experts reached 34% even with web access.
Humanity’s Last Exam (HLE)Hard academic questions; “HLE consists of 2,500 questions across dozens of subjects, including mathematics, humanities, and the natural sciences.”Each question “has a known solution that is unambiguous and easily verifiable, but cannot be quickly answered via internet retrieval.”
SWE-benchReal software problems: “Given a codebase along with a description of an issue to be resolved, a language model is tasked with editing the codebase to address the issue.”2,294 problems from 12 popular Python repositories; at launch the best model solved 1.96%.
SWE-bench VerifiedA smaller set of SWE-bench problems checked by people: “a subset of the original test set from SWE-bench, consisting of 500 samples verified to be non-problematic by our human annotators”Released by OpenAI with the SWE-bench authors in August 2024. See below for why OpenAI later stopped using it.
ARC-AGI-3Interactive puzzles: “There are no instructions, no rules, and no stated goals.”At its launch, the organisers reported: “Humans score 100%. Frontier AI scores 0.51%.”
METR time horizonHow long a task a model can handle: “the time humans typically take to complete tasks that AI models can complete with 50% success rate”METR’s paper reports this “has been doubling approximately every seven months since 2019, though the trend may have accelerated in 2024”. A January 2026 update estimated a doubling time of about 131 days for models since 2023.

Why a score can mislead

1. The test questions leak into training data

AI models are trained on huge amounts of text gathered from the internet, and benchmark questions are often published there too. Researchers call this data contamination: “the unintended overlap between training and test datasets”, which can lead to “an overestimation of the models’ true generalization capabilities.”

The clearest example comes from a benchmark’s own co-creator. In February 2026 OpenAI explained why it had dropped SWE-bench Verified. It said that “all frontier models we tested were able to reproduce the original, human-written bug fix used as the ground-truth reference, known as the gold patch, or verbatim problem statement specifics for certain tasks” — in other words, each model had seen at least some of the problems and their solutions during training. It concluded: “This is why we have stopped reporting SWE-bench Verified scores, and we recommend that other model developers do so too.”

2. The answer key can be wrong

The same OpenAI analysis found a second problem. It “audited a 27.6% subset of the dataset that models often failed to solve and found that at least 59.4% of the audited problems have flawed test cases that reject functionally correct submissions”.

That figure applies to the audited problems, not to the whole benchmark. But it shows how a model can be marked wrong for a correct answer — and why a score near the top of a test may say more about the test than about the model.

3. The way the test is run changes the result

OpenAI put it plainly in July 2026: “Benchmarks rarely measure AI models in isolation.” A model runs inside a harness — the software that feeds it the task, keeps track of its progress and collects its answer — and often at a chosen level of reasoning effort.

On 3 September 2026, ARC Prize published its results for OpenAI’s GPT-6 Astra on the Semi-Private test set of ARC-AGI-3. With its Standard harness the model scored 62.7% at maximum effort and 54.8% at high effort. With a Provider Adapter harness — which “preserves opaque reasoning state between requests and uses compaction for longer conversations, allowing the model to reuse prior work” — it scored 99.9% at high effort and 98.6% at maximum effort. The organisers’ summary put two of those runs side by side: “62.7% for $26K” and “99.9% for $19K”.

Same model, same benchmark, very different scores. A benchmark number without its setup is only half a fact.

4. The tasks may not be the tasks you care about

METR, which publishes the time-horizon measurements, lists the limits of its own benchmark. It says the time horizon is “a measure of the difficulty of a task, rather than the time an AI spends to complete the task.” It says: “Our task distribution is primarily composed of software engineering, machine learning, or cybersecurity tasks.” It warns: “In other words, our tasks are much “cleaner” than real economically valuable labor.” And its chart notes: “Measurements above 16 hrs are unreliable with our current task suite”.

A high score on a clean, well-specified software task does not show how a model will do on a messy task in a different field.

5. Who gets to test before release

Some leaderboards rank models by public votes rather than fixed questions. The best known is Arena (formerly Chatbot Arena, then LMArena), where people compare answers from two anonymous models and pick the better one. A 2025 paper, “The Leaderboard Illusion”, argued that its rules gave some developers an advantage. Its operators, then called LMArena, published a response disputing several of its claims.

Disputed The two sides agree that developers tested private versions of models before release. They disagree about whether that practice was undisclosed or favoured some developers, and about how much it distorted the rankings.

  • The paper reported that “undisclosed private testing practices benefit a handful of providers who are able to test multiple variants before public release and retract scores if desired”. It estimated that two large providers received 19.2% and 20.4% of all the arena’s data, while “a combined 83 open-weight models have only received an estimated 29.7% of the total data”, and reported “relative performance gains of up to 112% on ArenaHard, a test set from the arena distribution”, from training on arena data.
  • LMArena replied that its testing policy had been public since March 2024 and that “any model provider can submit as many public and private variants as they would like, as long as we have capacity for it”, while also saying that a future policy release would “explicitly state that model providers are all allowed to test multiple variants of their models pre-release”; that ArenaHard is “a static benchmark with 500 data points that uses an LLM judge, and no human labels”, and so not representative of the arena itself; and that its analysis put the effect of pre-release testing at “around +11 Elo after 50 tests and 3000 votes”, falling towards zero as new votes arrive.

Arena’s current policy, last updated in September 2026, says that if a model was tested before release, its score is marked “as preliminary until enough fresh votes have been collected after the model’s public release”.

The evidence does not settle these questions. The safe reading is that a public-vote ranking is one signal among several, not a final verdict.

6. A single “intelligence” number is a choice of ingredients

Some trackers combine many benchmarks into one index. Artificial Analysis says its “Intelligence Index v4.3 incorporates 10 evaluations”, and adds: “Like all evaluation metrics, it has limitations and may not apply directly to every use case.” Its methodology page now lists GPQA Diamond and MMLU-Pro under “Legacy Evaluations”.

Change the ingredients and the ranking can change. A combined score is useful for a quick comparison, but it reflects the tracker’s choice of which tests count.

The old rule behind all of this

The social anthropologist Marilyn Strathern wrote in 1997: “When a measure becomes a target, it ceases to be a good measure.” She drew on the work of an educationalist, Hoskin: “Hoskin describes this as ‘Goodhart’s law’, after the latter’s observation on instruments for monetary control”. Strathern was writing about exam grades in British universities, not AI — but the pattern is the same. Once a benchmark becomes the number everyone competes on, it becomes easier to improve the number than the ability it was meant to measure.

How to read a benchmark claim in two minutes

  1. Which benchmark, and which version? SWE-bench and SWE-bench Verified are different tests; so are ARC-AGI-2 and ARC-AGI-3.
  2. Who ran it? The model’s developer, the benchmark’s organisers, or an independent tracker. Results gathered by others are worth comparing: Epoch AI’s hub, for example, “includes results from benchmarks administered internally by Epoch AI as well as data collected from external sources.”
  3. With what setup? The harness and the reasoning effort can change the result.
  4. At what cost? Two scores reached at very different cost are not the same achievement.
  5. Is the test saturated? If every leading model scores near the top, small differences mean little.
  6. Could it be contaminated? OpenAI noted that SWE-bench Verified was “open-source and broadly used and discussed, which makes avoiding contamination difficult”. Widely published tests carry that risk.
  7. Are the tasks like yours? A coding benchmark tells you about coding, not about writing or research.

Our reading: benchmarks are the best shared measuring stick the field has, and they do track real progress. But a single score is a result under particular conditions, not a general measure of how capable a model is. Read the conditions before the number.

Sources

Checked September 2026.

Related: What is a frontier AI model? · Open-weight is not open-source · Numbers that mislead · How to read a study without a science degree

  • artificial intelligence
  • ai models
  • benchmarks
  • evaluation
  • explainer

SHARE & CITE

Hacks Vitae. "AI Benchmarks Explained: What the Scores Measure and Why They Can Mislead." September 17, 2026. https://www.hacksvitae.com/life-hack/ai-benchmarks-explained-what-the-scores-measure-and-why-they-can-mislead

That's what we found. The rest is your call.

118 articles, each with its sources listed. Spotted something off? hacksvitae@gmail.com

Open the library