Tests passing while complexity explodes from 29 to 285 is the perfect illustration of why benchmarks are misleading right now. The field keeps measuring "can AI write code" when the real question is "can it maintain software." Very different things.
AI Code Quality Beyond Tests: Complexity Metrics Matter
By
–