Actually it was (CC/codex/opencode) agents collaborating to *improve* Gemma 4
@thom_wolf
-
Multi-agent collaboration boosts Gemma 4 inference speed 5x
By
–
Multi-agents collaborations are among the most interesting agent behaviors right now!
— Thomas Wolf (@Thom_Wolf) 25 juin 2026
We did an experiment the other day with 100+ agents (an open-collaborations for a week) collaborating to improve the inference speed of Gemma 4 in vLLM. Got a 5x final improvement in speed but… pic.twitter.com/4PFS8L2mqeMulti-agents collaborations are among the most interesting agent behaviors right now! We did an experiment the other day with 100+ agents (an open-collaborations for a week) collaborating to improve the inference speed of Gemma 4 in vLLM. Got a 5x final improvement in speed but
-
Open-source models ensure humanity’s access to intelligence in AGI age
By
–
Open-source models will become a critical component of civilizational resilience in the AGI age. They will ensure that humanity retains access to a meaningful level of intelligence, regardless of the decisions of any individual actor.
-
AI moves to engineering artifacts: CADGenBench benchmark released
By
–
AI is moving beyond text, images, and code.
— Thomas Wolf (@Thom_Wolf) 8 juin 2026
Engineering artifacts are becoming a new class of model outputs and evaluating them requires different tools than we use for text, code, or images.
Today we're excited to release CADGenBench, a benchmark for CAD generation and… https://t.co/0JpyIbIdJTAI is moving beyond text, images, and code. Engineering artifacts are becoming a new class of model outputs and evaluating them requires different tools than we use for text, code, or images. Today we're excited to release CADGenBench, a benchmark for CAD generation and
-
Harvey Legal Benchmark: 1250 tasks across 24 practice areas
By
–
We’ll see it more and more Btw Harvey Legal Benchmark is very broad more than « super focused ». 1,250 legal tasks, 24 legal practice areas, it was made to be one of the first large-scale realistic legal benchmark
-
SDPO explored in practical async setups praised
By
–
cool work folks – nice to see SDPO explored in practical/async setups!
-

Extension to Terminal-Bench for Scientific AI Tool Evaluation
By
–
I'm very excited about this extension to the celebrated Terminal-Bench to science. If you're a scientist (life, physical, earth, mathematical science, etc) interested in AI, definitely check this out! Terminal bench evaluate how good AI models are at controling tools on a
-
The Rise of Vibe-Coded AI Dashboards as a Social Trend
By
–
My favorite 2026 trend is friends casually showing each other their vibe-coded life/work dashboards and AI setups. Ngl it feels exactly like bringing your Magic: The Gathering binder to school as a teen.
-
A new generation using AI to rebuild games
By
–
my 13 yo the other day: “we didn’t want to pay for the game with my friend so we just rebuilt it with Codex” me at 13 in the 90’s: HEX-editing the executable to NOP the license check because I didn't want to put all my pocket money in this game times they are a-changin’