To those in the replies who say 'but opus 4.8 is weaker so without fallback, the score would be even higher': this is not necessarily true because of how any benchmark works – which is an average of queries – and what is called 'the x.com/ClementDelangu…'
Debate on the effect of fallback in averaged benchmarks
By
–
