Yeah, not enough people understand how hard it is to do what the HF team does at scale (and they start complaining as soon as they're asked to pay $9 a month haha)
@clementdelangue
-
Avoiding value capture by a few dominant models
By
–
The last thing any of us wants is a world where every company in every sector gives up value to a few models that devour everything they see. If all value is captured by only a few models, the political economy simply does not
-

Two paths for AI: closed code or open-source
By
–
There is no inevitability in AI. We all have agency over what comes next: Path 1: Closed-code APIs, concentration of power, and a future decided by a handful of people in Silicon Valley and DC Path 2: Open-source AI, where everyone
-
Guardrails of cutting-edge APIs: superficial and bypassable
By
–
Many people have known for a while that the guardrails for cutting-edge model APIs are very easily bypassed, quite superficial, and impossible to fix. They are for the most part just a smokescreen and a distraction, in my opinion. We need a
-
Trip to DC to discuss open-source AI and transparency
By
–
I decided to go to DC next week to talk directly with policymakers. Not sure about the impact it will have, but with everything going on, it seems like a good time to say more about open-source AI, transparency, the concentration of
-

Debate on the effect of fallback in averaged benchmarks
By
–
To those in the replies who say 'but opus 4.8 is weaker so without fallback, the score would be even higher': this is not necessarily true because of how any benchmark works – which is an average of queries – and what is called 'the x.com/ClementDelangu…'
-

The absence of fallback does not necessarily increase the benchmark score
By
–
To those in the replies who say 'but opus 4.8 is weaker so without fallback, the score would be even higher': this is not necessarily true because of how any benchmark works – which is an average of queries – and what is called 'the x.com/ClementDelangu…'
-

The erroneous reasoning about fallback and benchmarks
By
–
To those in the replies who say "but opus 4.8 is weaker so without fallback, the score would be even higher": this is not necessarily true because of how any benchmark works – which is an average of queries – and what is called "the x.com/ClementDelangu…"
-
Refutation of the argument that a weaker model increases the score
By
–
To the people in the replies who say "but opus 4.8 is weaker, so without fallback, the score would be even higher": this is not necessarily true because of how any benchmark works – which is an average of queries – and what is called "the x.com/ClementDelangu…"
-
Benchmarks: why a weaker model does not necessarily improve the score
By
–
To those in the replies who say 'but opus 4.8 is weaker so without fallback, the score would be even higher': this is not necessarily true because of how any benchmark works – which is an average of queries – and what is called 'the x.com/ClementDelangu…'