A new benchmark just exposed the dirty secret behind every coding agent. Millions of developers now let AI agents write entire features unsupervised. A Carnegie Mellon paper tested whether that code is safe to ship. The team built SusVibes, a benchmark of 200 real coding
Carnegie Mellon benchmark exposes safety risks of AI coding agents
By
–
