Given that we saw OpenAI test better checkpoints on the arena, it does also seem like they have also optimised for cost / mass appeal rather than pushed the capabilities as far as they could. We also don't have almost any benchmarks for GPT-5-Pro, which is genuinely superb and
OpenAI’s GPT-5-Pro: Cost Optimization Over Capability Push
By
–