To the people in the replies who say "but opus 4.8 is weaker so without fallback, the score would be even higher": this is not necessarily true because of how any benchmark works – which is an average of queries – and what is called "the x.com/ClementDelangu…"
RESEARCH
-
Refutation of the argument on fallback in benchmarks
By
–
To those in the replies who say 'but opus 4.8 is weaker so without fallback, the score would be even higher': this is not necessarily true because of how any benchmark works – which is an average of queries – and what is called 'the'
-
AI diagnosis needs prospective rigorous assessment for real-world tasks
By
–
Yes. Widely relying on this for patient diagnosis and management w/o prospective, rigorous assessment for real world tasks until now.
As the authors point out: "scale, alignment and cross-domain reasoning may outweigh domain-specific
tuning as determinants of medical competency -
SambaNova congratulates MiniMax on M3 open-weight model launch
By
–
Congrats to our partners at @MiniMax_AI on the launch of MiniMax M3. Open-weight models continue to push the ecosystem forward, and we're excited to bring M3 to RDUs down the road. Looking forward to following what's built with it.
-
Progress in fine-grained 3D motion control for AI video
By
–
Fine-grained 3D motion control in AI video just got a little bit closer https://t.co/Uqi4lVJunR
— fofr (@fofrAI) 12 juin 2026Fine-grained 3D motion control in AI-generated video just got a little closer
-
Fable and Mythos for data work, not huge dataset loops
By
–
to clarify, looping Fable over huge datasets wouldn't be a good use of compute for many reasons but much of what goes into better models is DATA WORK: eval design, rubrics, error analysis, repairing noisy data Fable can do this!
and Mythos will do this too -
Data work key to model improvement: evals, rubrics, error analysis
By
–
obviously, but u are missing the point i'm saying most of what makes models better is DATA WORK: evals, rubrics, error analysis, and so on Fable can do this, Mythos will do this, future models will do this
-

Kimi K2.7 Code Proves Less Thinking Improves Coding Performance
By
–


The less your model thinks, the better it codes. Kimi K2.7 Code proved it. The entire reasoning model space has been moving in one direction. More thinking tokens, longer chains, bigger reasoning budgets. Kimi just challenged that. K2.7 Code scores higher than K2.6 on every
-

New protocol enables AI agents to evolve and improve themselves
By
–
What if AI agents could evolve themselves to get better over time? Researchers from NTU, Stanford, Princeton, and others introduce Autogenesis Protocol (AGP). It separates what agents use (prompts, tools, memory) from how they improve—letting agents track versions, propose
-
Olivier Rimmel congratulates Mistral and calls for three AI providers in France
By
–
The French AI #Mistral is very advanced and very performant, not at the level of Opus, even less of Fable, but it's still very good and it's improving all the time. I want three providers like that in France. President @EmmanuelMacron, make sure there are