you look at the diffs and see that the model improved structure, reduced repetition, broke apart large files etc.
MACHINE LEARNING
-

Reverse Jevons Paradox: Fable Could Reduce Spending
By
–
I wonder if Fable might cause a sort of 'inverted Jevons paradox', where companies would realize how much tokenmaxxing would cost them with Fable, to the point that they would reduce their spending beyond what they would have done if only Opus were
-
Edge AI 2026: Technologies Enabling On-Device Generative AI
By
–
Edge #AI in 2026: The Technologies Making On-Device #GenerativeAI Possible
— Ronald van Loon (@Ronald_vanLoon) 12 juin 2026
via @WevolverApp#ArtificialIntelligence #MachineLearning #ML #DL pic.twitter.com/kE5um3nNy4Edge #AI in 2026: The Technologies Making On-Device #GenerativeAI Possible
via @WevolverApp #ArtificialIntelligence #MachineLearning #ML #DL -

First infrastructure benchmark for agentic AI
By
–
The first infrastructure benchmark for agentic AI has arrived. An AI agent chains tens to hundreds of AI model calls, using tools, gathering context, and iterating until the task is complete. Existing benchmarks
-
GPT is 10-20x more token and cost effective for similar outcome
By
–
IMO sth that is a bit overlooked but will become far more important in the future. GPT is 10-20x more token+cost effective for ~similar outcome.
-

Apple uses NVIDIA Blackwell B200s on Google Cloud for private inference
By
–

I had already wondered how Apple manages to perform inference at Google while simultaneously protecting their privacy, essentially their unique selling point. The answer: the heaviest requests run on Blackwell B200s inside Google Cloud, with NVIDIA's Confidential Computing
-
Humanity’s final exam and FrontierMath incoming
By
–
Humanitys very last exam and Very FrontierMath incoming
-

Fable 5’s lead until GPT-5.6 and benchmark saturation
By
–

Looking at the graph, I think Fable 5 will only maintain its lead up to GPT-5.6. And secondly, I think the benchmark will soon be completely saturated.
-

Debate on the effect of fallback in averaged benchmarks
By
–
To those in the replies who say 'but opus 4.8 is weaker so without fallback, the score would be even higher': this is not necessarily true because of how any benchmark works – which is an average of queries – and what is called 'the x.com/ClementDelangu…'
-

The absence of fallback does not necessarily increase the benchmark score
By
–
To those in the replies who say 'but opus 4.8 is weaker so without fallback, the score would be even higher': this is not necessarily true because of how any benchmark works – which is an average of queries – and what is called 'the x.com/ClementDelangu…'