To those in the replies who say "but opus 4.8 is weaker so without fallback, the score would be even higher": this is not necessarily true because of how any benchmark works – which is an average of queries – and what is called "the x.com/ClementDelangu…"
AI
-
Refutation of the argument that a weaker model increases the score
By
–
To the people in the replies who say "but opus 4.8 is weaker, so without fallback, the score would be even higher": this is not necessarily true because of how any benchmark works – which is an average of queries – and what is called "the x.com/ClementDelangu…"
-
Benchmarks: why a weaker model does not necessarily improve the score
By
–
To those in the replies who say 'but opus 4.8 is weaker so without fallback, the score would be even higher': this is not necessarily true because of how any benchmark works – which is an average of queries – and what is called 'the x.com/ClementDelangu…'
-

Asynchronous AI Cuts Energy by Orders of Magnitude while Learning Continuously
By
–
Asynchronous #AI cuts computing energy by orders of magnitude while learning continuously
by Daegan Miller @TechXplore_com Learn more: https://
bit.ly/4fuihmz #MachineLearning #ArtificialIntelligence #DL #ML -
Debate on the impact of fallback in benchmarks
By
–
To the people in the replies who say 'but opus 4.8 is weaker so without fallback, the score would be even higher': that is not necessarily true because of how any benchmark works – which is an average of queries – and what is called 'the x.com/ClementDelangu…'
-
Refutation on benchmarks and the lack of fallback
By
–
To the people in the replies who say "but opus 4.8 is weaker so without fallback, the score would be even higher": this is not necessarily true because of how any benchmark works – which is an average of queries – and what is called "the x.com/ClementDelangu…"
-
Refutation of the argument on fallback in benchmarks
By
–
To those in the replies who say 'but opus 4.8 is weaker so without fallback, the score would be even higher': this is not necessarily true because of how any benchmark works – which is an average of queries – and what is called 'the'
-
AI diagnosis needs prospective rigorous assessment for real-world tasks
By
–
Yes. Widely relying on this for patient diagnosis and management w/o prospective, rigorous assessment for real world tasks until now.
As the authors point out: "scale, alignment and cross-domain reasoning may outweigh domain-specific
tuning as determinants of medical competency -
SambaNova congratulates MiniMax on M3 open-weight model launch
By
–
Congrats to our partners at @MiniMax_AI on the launch of MiniMax M3. Open-weight models continue to push the ecosystem forward, and we're excited to bring M3 to RDUs down the road. Looking forward to following what's built with it.
-
SambaNova excited about MiniMax M3 open-source and future RDUs
By
–
Excited to see @MiniMax_AI M3 now available to the open-source community. We're looking forward to bringing M3 to RDUs in the future. Congrats to the MiniMax team on the launch.
