I've long preferred Claude Code over Codex or Gemini, because it seemed much more reliable, but couldn't explain why : now Bullshit Bench by @petergostev provides compelling numbers. It measures bullshit as "when given false premises disguised in jargon, will the model go with
@aymericroucher
-
Lending stochastic parrots term, refining critique
By
–
Yeah, we’ve got to say it! (Well, after that, I did lend him the term “stochastic parrots,” which he doesn’t use, thanks @Fabien_Mikol
, I need to refine my critique) -

Me and my agents after 83 commits in a day
By
–
Me and my agents after I clocked 83 commits in a day
-
Qwen AI dominance and Alibaba leadership shakeup
By
–
Qwen team:
*each new model family obliterates SOTA*
*bonus vision*
*dense models are ngmi? Not on my watch*
*literally own the Pareto frontier for years*
*ultra based capybara mascot* Alibaba:
*nuke leadership* -

Smolagents creator discovers unauthorized Product Hunt launch
By
–
A similar thing happened to us with smolagents : within hours of publishing, someone congratulated me for the good launch on Product Hunt. I didn't do any. Turns out someone had created http://
smolagents.org and launched that for us. At first I feared a crypto scam, but -
Request to recreate an image of Thierry Breton showing a laptop
By
–
@grok redo this image but it's Thierry Breton showing off a laptop and saying "wow this is an impressive piece of innovation"
-
Computer Use moving from research to production with impressive demos
By
–
We're soon going to see Computer Use move from research to production use cases.
— m_ric (@AymericRoucher) 2 mars 2026
Standard Intelligence's demos are stunning, the model seems to have a very natural way of interacting with the computer. It can even drive a car through a webapp!
And Anthropic is working on it as… pic.twitter.com/FeGvZVFNeNWe're soon going to see Computer Use move from research to production use cases. Standard Intelligence's demos are stunning, the model seems to have a very natural way of interacting with the computer. It can even drive a car through a webapp! And Anthropic is working on it as
-
Reasoning requires basic knowledge for effective search guidance
By
–
You still need a bit of knowledge in any reasoning, just to guide you in where to look. For instance, you'll do better search on geopolitics if you know a bit about the general trends. Qwen just demonstrated that this base level of knowledge can be attained in a very compact
-

4B model rivals 120B on knowledge, ideal for agentic core
By
–


It's crazy that a 4B model can match GPT-OSS-120B **even on knowledge benchs** like MMLU and GPQA! Knowledge was long a limitation of smol models ; not anymore. This 4B could be a good candidate for the "agentic core" model described by @karpathy
-

Anthropic acquires Vercept, xAI focuses on computer use agents for 2026
By
–
– Anthropic acquires Vercept
– xAI's MacroHard heavily focuses on Computer use 2025 was the year of agents in CLI
2026 is shaping up to be the year of computer use agents.
