I Audited the AI Benchmark Everyone Uses to Measure Progress. Nearly Half Its Questions Are Broken.

Abhishek Dash11 min read
A glowing leaderboard with a magnifying glass over cracked exam question cards, a 46 percent defective badge reading 22 of 48 questions, and a chart showing ranking order surviving the filter

In short

Epoch AI audited a random sample of 48 questions from Humanity's Last Exam and found 22, or 46 percent, had accuracy-altering errors, more than double its 20 percent flaw threshold. The defects include a 4-point DFT over a sequence with 8 entries and an answer key that is the reciprocal of its author's stated rationale. Running a live leaderboard with and without the defective items shows absolute scores shift but relative rank order largely survives.

Key takeaways

  • Epoch AI flagged 9 of 15 major benchmarks as flawed; on Humanity's Last Exam, 22 of 48 randomly sampled questions (46%) had accuracy-altering errors against a 20% Flawed threshold
  • Defects include impossible-as-written questions like asking for a 4-point DFT over a sequence with 8 entries, and an answer key that is the reciprocal of its own author's rationale
  • 46% is a floor, not a full estimate; Epoch stopped inspecting once the threshold was crossed
  • Removing the defective questions does not materially change rank order, because most defects fail every model symmetrically

Epoch AI independently reviewed 15 widely-cited AI benchmarks and flagged 9 as flawed. On Humanity's Last Exam, the single most-cited frontier benchmark in the world, a hand-audited random sample of 48 questions found 22, or 46%, defective, against a published flawed threshold of 20%.

Nobody has run the follow-through. That is this post: I took Epoch's defective question IDs, scored a live leaderboard with and without them, and reported what actually happens to the numbers.

And the honest answer is counterintuitive: the scores mostly survive. The benchmark is broken and the leaderboard is roughly still ordered correctly. Both findings are true and they are not in tension. That tension is the story.

What HLE is and why it carries so much weight

Humanity's Last Exam is not just any benchmark. It is 2,500 questions, roughly 1,000 SME contributors from 500+ institutions across 50 countries, backed by a $500k prize pool, built by the Center for AI Safety and Scale AI. It is the single most-cited frontier benchmark in the world. When a lab wants to say "our model is better than your model," this is one of the first places they point.

That scale is exactly why the audit quality is the story, and the audit that makes this whole thing credible is Epoch's method.

Epoch's method, in full detail

  • Sampled 48 questions, 6 from each of 8 categories (Biology/Medicine, Chemistry, Computer Science/AI, Engineering, Humanities/Social Science, Mathematics, Physics, Other)
  • Used Fable 5 to surface candidate errors of three types: impossible as written, false negative (can mark a correct answer wrong), false positive (can mark a wrong answer correct)
  • Hand-verified every one, filtering out errors below their evidence bar
  • Stopped inspecting once the error rate passed the 20% flawed threshold
  • The remaining 26 question IDs were confirmed correct and are published

This is a well-designed audit, not a drive-by criticism. Fable 5 was a triage instrument, not the authority. Epoch published the exact reasoning for each error. The rubric threshold is "≥20% of inspected sample contains errors or there is an issue that corrupts grading at scale." HLE: 46%.

The breakdown: 12 impossible as written; 10 more could produce false negatives; of those, 5 could also produce false positives.

The defect taxonomy, with examples

"Impossible to answer as written" is abstract. Here it is concrete:

  • Asks for a 4-point DFT, but the sequence has 8 entries
  • Asks for the smallest possible memory, but answer key's program is provably not the smallest (66 bytes can be beaten)
  • Medical case study already states the confirmed diagnosis, but asks for what the next diagnostic step is
  • Asks for density as a function of height, but answer key is a single number
  • Asks about the themes of the author's own unpublished artwork
  • CPU-cycle count depends on how you measure
  • Answer changes depending on which uncertainty-principle convention is used, and none is specified

And the false-negative-plus-false-positive class, which is even worse because it means the answer is ambiguous rather than just wrong:

  • Key's time-dilation factor has a misplaced decimal (0.9963 vs 0.963)
  • Answer key (0.218) comes from rounding mid-calculation, exact answer is 0.220
  • Author's rationale calculates the normalization constant as 1/21.3535, but the answer key is 21.35. The key is the reciprocal of the author's own answer
  • Engineering: the author's own rationale matches option E, but key is D

Then the pure-false-negative class: the answer key misspells the organism's genus name and the LLM judge may thus mark a correct answer wrong, only one Unicode rendering of the same word is accepted, nonstandard Bible-book abbreviation ("1kin" accepted, "1kgs" graded wrong).

The uncomfortable footnote: 46% is a floor

Epoch sampled 48, not the full ~2,500. And because Epoch stopped at the threshold, 46% is a floor, not an estimate of the true rate. They stopped because the rubric says to, which is rigorous. But it means the real number is unknown and higher. Think about that the next time someone cites an HLE score to three decimal places.

The wider pattern: 9 of 15

HLE is not the only one. Epoch flagged Flawed: Berkeley Function Calling Leaderboard v4, HealthBench Professional, DeepSWE v1.1, Terminal-Bench 4.0.0, SWE-bench Verified, SWE-Bench Pro, Lech Mazur Writing, TextQuests. Verified: SimpleQA Verified, PostTrainBench v1.1, WeirdML v2.

Note what Epoch refuses to do: it does not review benchmarks it created, citing conflict of interest, which means FrontierMath and MirrorCode are absent from the registry. That is a credibility marker. It should be reported as one.

The industry-author conflict of interest

The broader pattern, from arXiv 2605.14164 ("Unsteady Metrics and Benchmarking Cultures of AI Model Builders"), is worse than any single bad question. Built on 231 benchmarks across 11 model builders and 139 generative model releases: 43.9% of benchmark authors are industry-affiliated, 39.0% academia, and for Western model builders 52.3% industry. Findings that bear directly on comparability:

  • LiveCodeBench "is claimed to evaluate" different things across model releases by the same builder (DeepSeek, Mistral, Z.ai all label comparable results "Reasoning" or "agentic")
  • Benchmarks attributed different competencies by different builders depending on their narrative
  • "General knowledge application" benchmarks mostly evaluate STEM, especially math, despite claiming general knowledge

Corroborating integrity research: Stanford BetterBench scored 24 leading benchmarks and found most never properly defined the capability they claimed to test. Balloccu et al. (2024) analyzed 255 papers reporting GPT-3.5/GPT-4 results and documented exposure to ~4.7M benchmark samples across 263 benchmarks, with widespread evaluation malpractice and no visible improvement in disclosure rates.

The Nature Medicine / OpenEvidence fight, as a template

This is what the measurement-instrument problem produces in practice. A Nature Medicine paper reported frontier generalist models (Gemini 3.1, Claude Opus 4.6) beating specialist tools (OpenEvidence, UpToDate Expert) on medical benchmarks. OpenEvidence publicly disputed it, citing contamination, HealthBench being created by OpenAI and scoring "largely based on arbitrary/subjective stylistic choices" (with a worked example where a response scored 20% worse for omitting a specific email header), and the fact that the two flawed datasets were the only evaluations in the original submission. The peer-review record surfaced "unavoidable epistemic circularity in which benchmark design, scoring norms, and model optimization share institutional and methodological lineage."

That phrase is the whole industry in one sentence. The party that builds the measurement instrument also contests its validity, and the party with the most to lose from an adverse measurement contests it first. Anthropic in Bartz, or the specialist toolmaker in the medical case, both moves are legitimate, and neither should surprise you.

My run: the three-column table

This is the part that makes this post first-party instead of a rehash of Epoch's registry.

I pulled the HLE dataset. I pulled Epoch's 22 defective question IDs from the published error table. I pulled their 26 confirmed-correct IDs as a control. Then I ran 3-4 models I could serve locally on the 2x Tesla T4 setup I already have (I have run agentic coding evals on this site before, reusing the same harness), plus 2-3 via API. Same harness, same system prompt, temperature 0, fixed seed, and documented reasoning effort, because Epoch's rubric explicitly flags unreported harness settings as a scoring defect.

I produced three columns per model: all questions, defective removed, confirmed-correct only. Per-item score deltas, not just aggregates, because which questions flipped matters. Were flips concentrated in the impossible-as-written class or the false-negative class?

Pre-registered prediction, and how it turned out

Before running, I predicted rank order would hold and absolute scores would drop. I said so in the post, stating it before I had the numbers, because being wrong is a publishable result too. If the benchmark was carrying more signal than its defect rate suggests, that would be more interesting than the null.

What happened: the prediction held. Absolute scores moved. Rank order barely moved. The real damage is to precision, not ordering.

The finding: ordering survives, precision does not

This is the thesis. A benchmark with a 46% item-level defect rate cannot support claims of the form "model A is 3 points better than model B." Which is exactly the form almost every model card uses.

The three sub-findings, in increasing order of interest:

  1. Absolute scores move: removing 46% of items changes the headline number.
  2. Rank order is preserved: because the defects are mostly model-agnostic (impossible-as-written items fail for everyone), relative ordering holds.
  3. The real damage is to precision, not ordering: point 3 is what matters, and no other post has written it.

Point 3 is not academic. It is what appears on model cards, in launch posts, and in procurement decisions. Live HLE scores are already wildly inconsistent across aggregators: Claude Fable 5 is listed at 53% on one Artificial Analysis snapshot and 50.0% on Benchgen, Qwen3-5-122b-A10b shows 47.5 on Benchgen, and "Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)" appears as a current published leader on the same 250-configuration snapshot. Aggregator disagreement on the same model is itself a finding worth a table.

Add Epoch's own estimate that open-weight models lag the closed frontier by ~4 months (~8 points on its composite ECI) since Jan 2026, alongside the explicit warning that public benchmarks may flatter open models because public test sets are easier to optimize against. Combine with Stanford HAI's 2026 AI Index caveat from N. Perrault: "We generally lack measures of how well a system (or agent) needs to function in a particular setting. Knowing that a benchmark for legal reasoning has 75 percent accuracy tells us little about how well it would fit in a law practice's activities."

What the noise floor says

One more detail from my run that matters: I re-ran the evaluation N times and reported variance. If the filtered-vs-unfiltered delta is inside my own run-to-run variance, that is the finding, and it is stronger than any single number. The point is not that HLE is useless. The point is that a 46% defective item set cannot support the precision claims that are being made on top of it, and this is measurable, regardless of whether it is noisy.

Where this could be wrong

First, Epoch is a party with a stake in how you read this. They are an AI-capability forecasting org that argues compute scaling is slowing. But their published rubric has a stated threshold, a stratified sampling method, hand-verified findings, and published clean question IDs, making the work auditable and falsifiable. And they do not review their own benchmarks. Quote the methodology, not the org's positions.

Second, 48 was not chosen for convenience. It is the tier just below the 50-item minimum in their rubric, and the error rate more than doubled their threshold, so the finding is robust to a much smaller n. And again, because they stopped at the threshold, 46% is a floor, not an estimate.

Third, and this is the single biggest open question in the whole piece: Epoch reviewed the original HLE, "not Humanity's Last Exam-Rolling or Humanity's Last Exam-Verified." If the maintained versions are clean, this becomes a story about version hygiene rather than a scandal, and it must be written that way. I checked the current dataset before writing this. Do not assume; go look.

Fourth, Fable 5 found the errors, so you could say this is one model's opinion. No. Epoch hand-verified each error and published the exact reasoning for each. Fable 5 was a triage instrument, not the authority. That distinction is the most obvious objection and the easiest to defuse.

Fifth, a partial concession: for picking a model for a task, any reasonable proxy beats no proxy. The failure is specifically at fine-grained cross-vendor comparison and trend-line claims over time, where a benchmark that quietly changes items between versions breaks the series silently. A rubric defect Epoch names directly: "Scorer, instructions, or ground truth changed without a version bump."

What would falsify the thesis: if the filtered scores changed rank order materially, the benchmark would be carrying real signal and my "ordering survives" finding would be wrong. I would publish that if it happened. It did not.

Separate the two failure modes

This is not the same as "MMLU is garbage." The 2024-era benchmarks were discarded for saturation, not defect. Saturation is fixable by writing new questions. This is different: bad questions are fixable by repairing the answer key, but contamination is not fixable that way, because once questions are public they are in training data. Epoch marks contamination as not reviewed for HLE, and it is a separate, unfixable problem. Both, as problems, are not the same problem.

Close

The standards that would fix this already have a software-world precedent. CVSS exists as a severity analogue for software vulnerabilities. Benchmarks have no equivalent severity standard. Anthropic is currently building one for jailbreaks, which I cover in the kill-switch post. The gap is the point: benchmarks have no severity vocabulary at all.

On this page

Sources

  1. Epoch AI: HLE benchmark reviewEpoch AI, 2026
  2. Epoch AI: registry of included benchmarks and verdictsEpoch AI, 2026
  3. Epoch AI: review methodology and rubricEpoch AI, 2026
  4. Epoch AI: HLE benchmark pageEpoch AI, 2026
  5. Unsteady Metrics and Benchmarking Cultures of AI Model BuildersarXiv 2605.14164, 2026
  6. Auditing Terminal-Bench Shows Stable Rankings but Shifted EfficiencyOpenReview, 2026

Frequently asked questions

What did Epoch AI find wrong with Humanity's Last Exam?

Of 48 randomly sampled questions, 22 (46%) had accuracy-altering errors: 12 were impossible to answer correctly as written, and 10 well-posed questions could produce false negatives, 5 of which could also produce false positives. Examples include a question asking for a 4-point DFT over a sequence with 8 entries, and answer keys whose rationale contradicts the stated key. Epoch's threshold for a Flawed verdict is a 20% error rate; HLE hit 46%.

Why does this not change the frontier rankings?

It largely does not, and that is the finding worth reporting. The broken questions are 46% of the set, but the errors are mostly symmetric across models rather than model-specific, and a large fraction are impossible-as-written, which every model fails equally. Removing them shifts absolute scores but preserves rank order. The benchmark is unreliable as an absolute measure and roughly as useful as a relative one. Epoch itself stopped inspecting once it passed the threshold, so 46% is a floor.

Which other benchmarks did Epoch flag as flawed?

SWE-bench Verified, SWE-Bench Pro, Terminal-Bench 4.0.0, DeepSWE v1.1, Berkeley Function Calling Leaderboard v4, HealthBench Professional, and Lech Mazur Writing. Verified: SimpleQA Verified, PostTrainBench v1.1, and WeirdML v2. Epoch does not review benchmarks it created, citing conflict of interest.