Researchers have found that major AI agent benchmarks can be exploited to achieve near-perfect scores without actually solving tasks, by manipulating the evaluation environment or exploiting vulnerabilities in the benchmarks. The researchers' automated scanning agent was able to achieve high scores on eight prominent benchmarks, including SWE-bench, Terminal-Bench, and WebArena, without using any reasoning or capability. The exploits range from simple to complex and highlight the need for more robust and secure benchmarking methods. The findings suggest that the current benchmarking system is flawed and can be gamed, which can lead to misleading results and overestimation of AI capabilities.