That Haiku number is the one worth sitting with. Claude Haiku scores 39% on SWE-Bench Pro. On DeepSWE, where it can't coast on contaminated data or exploit the test environment, it scores zero. Not low. Zero. That's not a model that dropped in performance. That's a model that
Claude benchmarked: SWE-Bench Pro vs DeepSWE performance
By
–