So GPT-5 is running on GB200, to take full advantage of it, it could be in FP4. So I wonder if this change is hiding the size of the model somewhat. If it was FP8 and on H100, it would be now running at something like 10 tokens per second rather than current 50 tps.
LLMS
-

Google Releases Gemma3 270M Tiny LLM on KerasHub
By
–
Google just dropped a new tiny LLM with outstanding performance — Gemma3 270M. Now available on KerasHub. Try the new presets `gemma3_270m` and `gemma3_instruct_270m`!
-
OpenAI’s GPT-5-Pro: Cost Optimization Over Capability Push
By
–
Given that we saw OpenAI test better checkpoints on the arena, it does also seem like they have also optimised for cost / mass appeal rather than pushed the capabilities as far as they could. We also don't have almost any benchmarks for GPT-5-Pro, which is genuinely superb and
-
LLM Tool Organization: Role-Based Access Control Solution
By
–
I think we need more organization principles around what tools are given to the LLM/AI. Ideally, organizing it by roles, purpose and even layering on access control. We're working on a solution for this at MintMCP
-
Super-Efficient Tiny AI Model Released
By
–
It's a super-efficient tiny model. You can learn more about it here
-
LFM2-350M Model Release Announcement Oversight
By
–
Congrats on the release, but you forgot LFM2-350M
-
LangChain Academy Launches Deep Research with LangGraph Course
By
–
🔥 Our latest LangChain Academy course – Deep Research with LangGraph – is now live! 🔥
— LangChain (@LangChain) 14 août 2025
Deep research agents are taking off – from major AI labs to companies building their own.
Research is inherently open-ended. You can't always predict whether a question needs broad… pic.twitter.com/BDyNq6zlxcOur latest LangChain Academy course – Deep Research with LangGraph – is now live! Deep research agents are taking off – from major AI labs to companies building their own. Research is inherently open-ended. You can't always predict whether a question needs broad
-

Mol-R1: Explicit Long-CoT Reasoning for Molecule Discovery
By
–
Mol-R1 Towards Explicit Long-CoT Reasoning in Molecule Discovery
-
GPT-5 Pro API vs Chat Performance Gap Reassessed
By
–
Aaah intéressant ! J’ai eu de très nombreux retours indiquant que GPT-5 Pro via API était largement boosté par rapport au chat. Ils ont peut-être rééquilibré les deux, ce qui serait une bonne chose. En tout cas, en recherche théorique, il est très, très bon, sans doute le
