10 agent benchmarks. 219 vulnerabilities. Exploits that score near-perfect without solving the actual task. When an agent gets a high score, did it really do the work — or did it game the benchmark? Dartmouth College, UC Berkeley, and BenchFlow AI present BenchShield. A prior
BenchShield exposes 219 vulnerabilities in AI agent benchmarks
By
–
