On ARC AGI 2, Opus is just a smarter model and I think we see such big climb not because it is so good at thinking, but rather without thinking it is nearly impossible for it to do these puzzles well (same for me). E.g. GPT-5-Pro is below Opus 4.5 thinking 16k, but there's no way
Opus outperforms GPT-5-Pro on ARC AGI reasoning tasks
By
–