AI Dynamics

Global AI News Aggregator

About

Claude Opus 4.6 surpasses benchmark limits

The benchmark broke before the model did. METR's own task suite is saturated. They can't measure how capable Claude Opus 4.6 actually is because the tests aren't hard enough anymore. The headline number: 14.5 hours of autonomous software work at 50% success rate. But the real

→ View original post on X — @godofprompt