breaking: bad news for Anthropic since Meta was said to be a big customer and is cutting its token budgets. more generally lots of companies will make the same decision; next year’s token budgets won’t be the freewheeling affair they were earlier this spring. honeymoon is
LLMS
-
Humanity’s final exam and FrontierMath incoming
By
–
Humanitys very last exam and Very FrontierMath incoming
-

Fable 5’s lead until GPT-5.6 and benchmark saturation
By
–

Looking at the graph, I think Fable 5 will only maintain its lead up to GPT-5.6. And secondly, I think the benchmark will soon be completely saturated.
-

Debate on the effect of fallback in averaged benchmarks
By
–
To those in the replies who say 'but opus 4.8 is weaker so without fallback, the score would be even higher': this is not necessarily true because of how any benchmark works – which is an average of queries – and what is called 'the x.com/ClementDelangu…'
-

The absence of fallback does not necessarily increase the benchmark score
By
–
To those in the replies who say 'but opus 4.8 is weaker so without fallback, the score would be even higher': this is not necessarily true because of how any benchmark works – which is an average of queries – and what is called 'the x.com/ClementDelangu…'
-

The erroneous reasoning about fallback and benchmarks
By
–
To those in the replies who say "but opus 4.8 is weaker so without fallback, the score would be even higher": this is not necessarily true because of how any benchmark works – which is an average of queries – and what is called "the x.com/ClementDelangu…"
-
Refutation of the argument that a weaker model increases the score
By
–
To the people in the replies who say "but opus 4.8 is weaker, so without fallback, the score would be even higher": this is not necessarily true because of how any benchmark works – which is an average of queries – and what is called "the x.com/ClementDelangu…"
-
Benchmarks: why a weaker model does not necessarily improve the score
By
–
To those in the replies who say 'but opus 4.8 is weaker so without fallback, the score would be even higher': this is not necessarily true because of how any benchmark works – which is an average of queries – and what is called 'the x.com/ClementDelangu…'
-
Debate on the impact of fallback in benchmarks
By
–
To the people in the replies who say 'but opus 4.8 is weaker so without fallback, the score would be even higher': that is not necessarily true because of how any benchmark works – which is an average of queries – and what is called 'the x.com/ClementDelangu…'
-
Refutation on benchmarks and the lack of fallback
By
–
To the people in the replies who say "but opus 4.8 is weaker so without fallback, the score would be even higher": this is not necessarily true because of how any benchmark works – which is an average of queries – and what is called "the x.com/ClementDelangu…"
