AI Dynamics

Global AI News Aggregator

About

New Paper Evaluates AI Coding Agents on Long-Horizon Codebase Maintenance

Even the SoTA models still struggle with long-horizon code maintenance! “SWE-CI: Evaluating Agent Capabilities in Maintaining Codebases via Continuous Integration” Right now, most coding-agent evals are static, where you solve a single issue, and pass one-off tests. This paper

→ View original post on X — @askalphaxiv