On the 10 benchmarks where both OpenAI and Anthropic report scores, here's the split: Claude Opus 4.7 leads on 6. GPT-5.5 leads on 4. But the leads aren't random. They cluster by category. And that changes what "winning" means entirely.
Claude Opus 4.7 leads GPT-5.5 on 6 of 10 benchmarks by category
By
–
