that’s not true, evals comparing 5.2 vs 5.4 (thinking only) Coding (SWE-Bench Pro) > GPT-5.2: 55.6%
> GPT-5.4: 57.7%
→ ~+2.1 pts improvement in solving real-world repo bug-fix tasks. Knowledge-work benchmark (GDPval) > GPT-5.2: ~71% win/tie vs professionals
> GPT-5.4: 83%
GPT-5.4 outperforms GPT-5.2 on code and knowledge benchmarks
By
–