To the people in the replies who say 'but opus 4.8 is weaker so without fallback, the score would be even higher': that is not necessarily true because of how any benchmark works – which is an average of queries – and what is called 'the x.com/ClementDelangu…'
@clementdelangue
-
Refutation on benchmarks and the lack of fallback
By
–
To the people in the replies who say "but opus 4.8 is weaker so without fallback, the score would be even higher": this is not necessarily true because of how any benchmark works – which is an average of queries – and what is called "the x.com/ClementDelangu…"
-
Refutation of the argument on fallback in benchmarks
By
–
To those in the replies who say 'but opus 4.8 is weaker so without fallback, the score would be even higher': this is not necessarily true because of how any benchmark works – which is an average of queries – and what is called 'the'
-
Proposal to evaluate Omni and build a routing system
By
–
Yes maybe we should evaluate Omni @victormustar or https://openrouter.ai/docs/guides/routing/provider-selection … @alexatallah? Or does someone want to build a routing system and evaluate on AA?
-

The graph illustrates the bias of AI evaluations toward closed APIs
By
–
This graph illustrates what is flawed in AI evaluations: they structurally favor closed APIs that can route, switch, ensemble and optimize behind the scenes without any transparency. No offense, @ArtificialAnlys, but how to compare a
-

@clementdelangue — 2026-06-11
By
–
HF est devenu la meilleure plateforme de stockage pour les modèles et les jeux de données PRIVÉS et PUBLICS, qu'ils soient intermédiaires ou finaux ! Excellent exemple de @heyjasperai qui a utilisé les buckets HF pour stocker leur jeu de données Monet et entraîner des modèles
-

AI must avoid any manipulation, even well-intentioned
By
–
Thank you, that's much better! Manipulation by AI must be avoided at all costs, even when it is well-intentioned!
-
Training an open source AI model for construction?
By
–
Should we try to train an open source AI model for construction? We obviously have interesting datasets with HF, MLintern, transformers, trl…
-
Gemma Challenge: Google and Hugging Face for Open-Source AI
By
–
Announcing the Gemma challenge!
— clem 🤗 (@ClementDelangue) 10 juin 2026
Google, Hugging Face, and the open-source AI community choose to empower AI builders rather than sabotage them.
Fun to see the Hub becoming the platform where agents collaborate, just as it became the platform where humans collaborate.… https://t.co/b8Kd6kPCWA pic.twitter.com/FeQKqEz2htLet's announce the Gemma Challenge! Google, Hugging Face, and the open-source AI community choose to empower AI creators rather than sabotage them. Fun to see the Hub become the platform where agents collaborate, just as it became the platform where
