Yes I know. We helped show the PyTorch team how to train Llama efficiently – they have a whole GPU Mode talk about it. These issues do not at all line with your claims however.
GENERATIVE AI
-
Software Eating World While AI Digests Technology
By
–
Software is eating the world, and AI is digesting it.
-

Why Bigger Isn’t Always Better in AI: LLMs vs SLMs
By
–
💡 Why Bigger Isn’t Always Better in AI
— Dr. Debashis Dutta (@debashis_dutta) 5 août 2025
🔎 Rethinking LLMs vs. SLMs — Through a Business Lens
In today’s GenAI landscape, organizations often default to Large Language Models (LLMs) — powerful, yes, but not always practical.
🔍 The reality?
Most business use cases don't require… pic.twitter.com/zpJ5OAdjvhWhy Bigger Isn’t Always Better in AI Rethinking LLMs vs. SLMs — Through a Business Lens
In today’s GenAI landscape, organizations often default to Large Language Models (LLMs) — powerful, yes, but not always practical. The reality?
Most business use cases don't require -

OpenAI’s Open-Source Models Beat Commercial Pricing Benchmarks
By
–
As benchmarks for OpenAI's new open-source models roll in, one thing to keep in mind is how cheap they actually are: – GPT OSS 20b is cheaper than Gemini 2.5 Flash Light or GPT-4.1-Nano.
– GPT OSS 120b is cheaper than recent open-source models from China (e.g., Kimi K2 or GLM -

GPT-OSS Reaches Number 1 Trending on Hugging Face
By
–
We’re already #1 trending on @huggingface with gpt-oss! Huge thanks to the open-source AI community for the love.
-

US Federal Agencies Gain Secure Claude Access for Operations
By
–
U.S. federal departments and agencies can now more quickly and easily get access to Claude to transform how they work, all while still meeting federal security and compliance requirements.
-
Tool Calling Implementation Challenges in New Harmony Models
By
–
Sounds to me like it's tool calling, which is the bit of the model I've not got working yet myself – the new Harmony stuff means there are plenty of details that model serving tools need to nail down before I'll be confident if we know how well that actually works
-
Future SOTA AI Models Training on Massive Token Sequences
By
–
clearly in five years SOTA AI models will train on a single string containing approximately 2^21 tokens
-

Open Source AI Model Training Details and Documentation
By
–
Added an extra section with interesting details from the model card about how the models were trained https://
simonwillison.net/2025/Aug/5/gpt
-oss/#the-model-card
… -
OpenAI Releases Two Open Weight Models Including gpt-oss-120b
By
–
I'm thrilled @OpenAI has released two open weight models. Thank you to all my friends at OpenAI for this gift! I'm also encouraged that from my quick tests gpt-oss-120b looks strong (though we should still wait for rigorous 3rd party evals).