AI Dynamics

Global AI News Aggregator

About

Claude Opus 4.5 hits 80.9% on SWE-bench, first to break 80%

the benchmark data made me pause. claude opus 4.5 hit 80.9% on SWE-bench verified. first model to ever break 80%. this isn't leetcode. it's real github issues from production repos. the actual work developers do. 4 out of 5 real-world bugs. solved.

→ View original post on X — @godofprompt