The benchmark broke before the model did. METR's own task suite is saturated. They can't measure how capable Claude Opus 4.6 actually is because the tests aren't hard enough anymore. The headline number: 14.5 hours of autonomous software work at 50% success rate. But the real
Claude Opus 4.6 surpasses benchmark limits
By
–
