I would be very interested to understand the difference in impact between training and inference, for example – these numbers bundle those together
@simonw
-

AI Lab Details Environmental Impact Numbers Analysis
By
–
First time I've seen a large AI lab break down the environmental impact numbers in this level of detail
-
Gemini Receives IMO Official Confirmation While OpenAI Awaits
By
–
It looks to me like Gemini may have got official confirmation from IMO while OpenAI did not
-
OpenAI Gemini Teams Recruit IMO Medal Winners
By
–
The OpenAI team had at least one previous IMO medal winner and I wouldn't be surprised if the larger Gemini team had some of those too
-

Gemini Solves IMO Problems: Comparing with DeepMind Achievement
By
–
Wrote this up for my blog, including a note comparing this achievement with the DeepMind result from IMO last year https://
simonwillison.net/2025/Jul/21/ge
mini-imo/
… -
Gemini and OpenAI tested models without tools or internet access
By
–
Same as OpenAI the Gemini team ran the model with no extra tools and no internet access
-
Gemini’s Tool Usage vs Model Inference Capabilities
By
–
Did that Gemini have tool usage (Python or Lean or similar) or did it solve the problems using model inference alone?
-
LLM Chat Apps Need Better Documentation Access Tools
By
–
I wish LLM chat apps would run a form of RAG for this – or just provide a tool called "help_answer_questions_about_abilities()" which, when called, dumps a few thousand extra tokens of documentation into the context Feels like low hanging fruit for a significant usability win
-
Grok 4 System Prompt Transparency on AI Capabilities
By
–
I find this really annoying, because asking questions about the system's capabilities is such a natural thing for people to do Grok 4's system prompt does at least have a whole section dedicated to what products xAI offer:
-
OpenAI and Gemini Tie at 35/42 on Challenge Benchmark
By
–
Interestingly, both OpenAI and Gemini achieve the exact same score: 35/42 – and both teams solved problems 1-5 but did not solve 6, the most challenging problem