Maybe we should come up with some more world model tests that could show it more definitively, one thing that is hard is whether we are just testing the language component of the model which confuses things. I haven't obviously noticed that GPT was worse at world model stuff, but
@petergostev
-
AGI Infographics Reliability and Precision Challenges
By
–
if we get the 'agi of infographics' I would maybe agree with you, but right now there's still just enough imprecision and sloppiness that means that they are just not reliably usable. If this is the dimension that people care about, then maybe Google will solve it
-

Nano Banana Pro vs GPT-Image-1.5: Image Generation Comparison
By
–
This topic attracts a lot of emotions, so we wanted to dig into where does Nano Banana Pro or GPT-Image-1.5 wins. Look for yourself, but this is my take: – Nano Banana (NB) is better at infographics
– NB avoids the 'AI look', while GPT images occasionally output very Dalle -
Impressive non-thinking mode performance in AI models
By
–
There's a thinking mode for it, but the difference isn't huge, the fact that the non thinking mode works so well is very impressive
-
Latest LLM Advances: Opus 4.5, GPT-5.2-xHigh, Gemini 3 Flash
By
–
Positive about where LLMs are going based on the last few weeks: – Amazing non-thinking model (Opus 4.5) – Impressive super long thinking model (GPT-5.2-xHigh) – Genuinely good smaller model (Gemini 3 Flash) Progress on all fronts!
-
GPT-1.5 vs Nano-Banana Pro: Performance and Prompt Engineering Comparison
By
–
My anecdotal impression of GPT-1.5 vs Nano-Banana Pro is that they are pretty neck & neck overall. I find GPT a lot easier to prompt, with NB, you often had to iterate several times before getting a good result, while with GPT you typically get what you ask for. But I think NB
-

GPT-Image-1.5 Capabilities: Exploring AI Image Generation Results
By
–
What can gpt-image-1.5 do? Had a lot of fun playing with it in the last few days. We've pulled together a compilation of some of our favourite results in this video. https://t.co/NeTgUPDOv7
— Peter Gostev (@petergostev) 16 décembre 2025What can gpt-image-1.5 do? Had a lot of fun playing with it in the last few days. We've pulled together a compilation of some of our favourite results in this video.
-
GPT and Claude struggle with self-improvement on visual tasks
By
–
Maybe gemini is better at this, but GPT and Claude models were pretty terrible at self-improving, if you give it a screenshot, its ability to actually pick up on obvious issues is pretty poor. I have tried this at the beginning, it came up with 10 different ways how to 'improve'
-
Lateral problem-solving: exploring root causes beyond surface symptoms
By
–
yeah nothing special, just asking it to fix stuff. occasionally if it was struggling I was trying to encourage it to explore the issue a bit more laterally. e.g. if the sky is not blue it doesn't mean you just need to crank up the colour, it is probably that something is blocking
-

AI System Iteration Issues and Technical Limitations
By
–
I have tried on the first day, but there was something wrong with it – it kept refusing to do any iterations, need to revisit in case they've fixed it