Did we ever learn what model won gold at the IMO from OpenAI? It was a year ago and it was called an unreleased internal general purpose model back then. Has GPT-5.5 Pro Extended caught up with whatever it was?
@emollick
-
Gemini 3.5 Flash miscounts letters — frontier still jagged
By
–
The frontier is still jagged though (here is Gemini 3.5 Flash messing up counting letters in words) https://
x.com/RRiscio37389/s
tatus/2057193260745883962?s=20
… -

Timeline: LLMs progress from simple errors to IMO gold and major theorem solve
By
–


June 2024: The latest general-purpose LLMs could not count the r's in strawberry.
July 2025: The latest general-purpose LLMs get gold in the International Math Olympiad.
May 2026: The latest general-purpose LLM solve one of the "best-known questions in combinatorial geometry" -
Recursive self-improvement concentrates AI talent and raises barriers to rivals
By
–
One interesting side feature of recursive self-improvement, to the extent that is happening, is that it makes the Big Three labs more appealing to talent, and shortens the runway for launching a potential competitor instead at the same time.
-
Google releasing AI tools to accelerate science
By
–
Got to play with a little of this before launch as well. My experience as a social scientist was that it was more bioscience focused right now, but I think Google has been the leading lab in releasing serious AI tools to accelerate science & expect to see them improve fast.
-
Major AI platforms converging or diverging: who will win?
By
–
The gap between what you can do on ChatGPT/Codex and Claude/Code/Cowork is closing, as Anthropic & OpenAI converge on a single experience. Google's experiences are diverging: Studio & Gemini & Antigravity & the other Google AI apps are increasingly different. Which will win?
-

Antigravity provides better end-of-task transparency than Codex
By
–
Fascinatingly Antigravity is actually the best tool so far at providing this sort of transparency, doing something by default that Codex and Code do not: offering a summary of exactly what it did at the end of a task. Just add this to Gemini! (But also cite sources more)
-

Antigravity reveals model thinking traces
By
–
The crazy thing is that Antigravity does this quite well! So it isn't like Google is hiding these thinking traces because of distillation or because they are full of insane mutterings. I guess you need to either do the work in Antigravity or not at all?
-

User critique: Gemini models not ready for enterprise
By
–

I find this continually frustrating. Gemini 3.5 Flash is excellent, as is Gemini 3.1 Pro. But you absolutely cannot use them for any serious purpose right now, especially for any enterprise work. Compare to Claude or ChatGPT: you can understand what the model did & how to correct
-
Critique of Gemini’s thinking traces and auditability
By
–
(I originally wrote that Gemini removed thinking traces. People in the comments pointed out that they are accessible from a menu, so I did a new tweet. But I cannot believe how useless the thinking traces are. It makes Gemini outputs completely unauditable and thus untrustworthy)