@eliebakouch could you do a comparison of Muon(Clip) vs AdamW to explain the differences?
@aymericroucher
-
Huggingface explains stateless direct response MCP server choice
By
–
If you're developing MCP servers, you should give a read to how the @huggingface team built the Hub MCP, they explain why they chose a Stateless + Direct Response server over other options!
-
SmolLM3: Powerful built-in tool-calling capabilities
By
–
Reminder: SmolLM3 comes with built-in tool-calling, and it works really well! pic.twitter.com/uba78UfPtg
— m_ric (@AymericRoucher) 9 juillet 2025Reminder: SmolLM3 comes with built-in tool-calling, and it works really well!
-

FlashAttention less useful with MLP GEMM latency dominance
By
–
Maybe FlashAttention is not that useful when you have MLP GEMMs that eat so much latency? Interesting graph in the latest blog post from @gpus_go_brrr!
-

SmolLM3-3B fills Qwen’s Pareto gap via agentic post-training
By
–
Qwen left a hole in the Pareto frontier of optimal performance for a given size… So we just filled it: introducing SmolLM3-3B I helped the SmolLM team on the "make it agentic" part, by post-training the model on agent traces with @akseljoonas
: the model is now also on -

Agents too unreliable, use only when no other choice
By
–
Agents are too unreliable => Use them only when you have no choice! > Don't get me wrong, agentic apps are powerful and unlocks vast fields of previously impossible use cases, but indeed people often try to use thel in uses cases where they don't belong. @hugobowne just
-
Stop Building AI Agents article by Hugo
By
–
Hugo's article: https://
decodingml.substack.com/p/stop-buildin
g-ai-agents
… -
Shocking Hub MCP cheat code supercharges IDE with semantic search
By
–
SHOCKING cheat code for your IDE: use the Hub MCP! This gives your Cursor/Claude/whatever access to tools that supercharge it for the Hub.
— m_ric (@AymericRoucher) 2 juillet 2025
The new Docs Semantic Search tool feels like intravenous caffeine supply to correct API errors in a few seconds, gg @mishig25 ⚡️⚡️
To… pic.twitter.com/083S79UUu3SHOCKING cheat code for your IDE: use the Hub MCP! This gives your Cursor/Claude/whatever access to tools that supercharge it for the Hub. The new Docs Semantic Search tool feels like intravenous caffeine supply to correct API errors in a few seconds, gg @mishig25 To
-

Can AI models self-improve recursively via agentic methods?
By
–
Could AI models start self-improving recursively, scaling up performance to the moon in an infinite loop? A friend said that it's impossible, I wasn't that sure. I argued that with agentic methods like Absolute Zero (basically an LLM generates training data by itself by
-

Kimi-Dev-72B shows impressive SWE-Bench Verified performance, no tech report yet
By
–
Why isn't anyone talking about Kimi-Dev-72B, released ~1 week ago on the Hub?
The tech report is not out yet, and I'll love to see more benchmarks, but performance on SWE-Bench Verified seems very impressive!