It wasn't just OpenAI who got gold on the International Mathematics Olympiad this year – here's Google Gemini's result
@simonw
-
Kimi K2 on Mac Studios: Premium AI Hardware Setup Cost
By
–
Closest you can get right now is probably Kimi K2 on a pair of Mac Studios, for about $20,000 https://t.co/9hzx5y3nmJ
— Simon Willison (@simonw) 20 juillet 2025Closest you can get right now is probably Kimi K2 on a pair of Mac Studios, for about $20,000
-
Defining AI Agents: OpenAI Anthropic Disagreement
By
–
As always, that depends very much on much on how you (and they) define "agents" When not even OpenAI and Anthropic can settle on a shared definition between the two of them there's no hope for anyone else!
-
LLM Capabilities Without Tools Exceed Expectations in Math
By
–
I thought that was true, but apparently I was wrong – LLMs without tools are a lot more capable than I expected, at least for things like IMO math problems I still think tools are the most important technique in AI engineering generally
-

Comparing Local LLM Sizes to Wikipedia Dumps
By
–
Evan Hahn made this handy table comparing the size of different LLMs to the size of Wikipedia dumps in various shapes and formats https://
evanhahn.com/local-llms-ver
sus-offline-wikipedia/
… -
MCP Architecture Flaw Discussion and Analysis
By
–
Yeah it's a pretty big flaw in the way MCP works, here's a good talk about that
-
Model Limitations with Multiple Tools vs Bash Python
By
–
Not exactly – for many models it appears that having more than about a dozen MCP or function-calling style tools can confuse the model, but you can successfully give them Bash or Python which is effectively thousands of tools all in one go
-
AI Achieves Task Without Tool Calls Innovation
By
–
I am not enough of a mathematician to have an interesting perspective on this, but I do think it is notable that they achieved this without tool calls – I had assumed that those would be necessary for this kind of task
-

Experimental Reasoning Model Achieves Score Without Tool Usage
By
–
The most notable thing about this result is that this unnamed experimental reasoning model achieved this score without any tool usage at all – it looks like it's just another classic next-token-predicting LLM with a bunch of reinforcement learning layered on top
-
Naming AI-Assisted Programming: From Vibe Coding to Standard Practice
By
–
We never found something snappy, which is why vibe coding grabbed the while gap I think it's just going to be "programming" pretty soon to be honest I call it "AI-assisted programming" but that's far too long and not catchy at all
