Claude Opus 4.1 has about 2% better evals in the SWE benchmark. Moreover, it is already available. But pay even more attention to the comment afterwards: In the next few weeks (!) much bigger improvements will be released for their models. We are accelerating!
Claude Opus 4.1 SWE Benchmark Improvement and Upcoming Model Updates
By
–
