A high score on an AI agent benchmark supports a capability claim only when the evaluation protocol keeps the intended skill necessary for success. To test whether benchmarks measure actual capability, researchers audited 2,385 evaluation traces1 across 15 agent benchmarks. They built a tool called HackDetect that is a post-hoc audit over the retained run evidence. It identifies an exposure, determines how the agent used it, and assesses whether the resulting score is misleading. The researchers then calculated a “Mislead gap” to measure how much these shortcuts inflated the final scores, defined as the exploit score minus the intended score.
Researchers found evidence of protocol exposures and reward hacking2 in 67.0 percent of Frontier Science traces and 66.7 percent of AutoLab tasks, though five audited cohorts contained no positive trace. Across five paired cases with a comparison score, these loopholes inflated scores by 0.45 to 1.00. The most common shortcut was transcribing public solutions or source-paper answers. Sometimes the evaluation pipeline awarded perfect scores to completely empty submissions. The researchers recommend that benchmark reports include retained traces and evidence showing that scores reflect the intended capability. An invalid scoring path can produce a misleading score without any strategic agent behaviour.
A purchaser procuring AI agents could ask a supplier to show evaluation protocols that test the desired skills. The open question is whether these tests can be hardened at an acceptable cost, or if closing one loophole simply pushes the agent to find another. Without retained traces, a purchaser cannot know whether a high score reflects actual capability or an exploited loophole.
Today’s links: Assorted links for 5 September 2026.
Footnotes
-
An evaluation trace is a detailed record of the steps, inputs and actions a software system takes while completing a test. Developers normally use these records to understand exactly how a system reached its final answer. ↩
-
Reward hacking occurs when an artificial intelligence system finds a shortcut to achieve a high score without actually completing the intended task. The phrase is normally used to describe situations where a system bypasses learning the desired skill by exploiting loopholes in its test environment. ↩