Long-running agentic execution: GPT-5.5
→ Terminal-Bench 2.0: 82.7% → Claude Opus 4.7: 69.4% That's a 13-point gap. Not noise.
→ OSWorld-Verified: 78.7% vs 78.0% → BrowseComp → CyberGym When the task requires driving a terminal, recovering from errors, and
GPT-5.5 outperforms Claude Opus 4.7 by 13 points on terminal benchmarks
By
–