Man, what's going on with this benchmark, it is literally getting worse with each release. "OpenAI-Proof Q&A evaluates AI models on 20 internal research and engineering bottlenecks encountered at OpenAI, each representing at least a one-day delay to a major project and in some
@petergostev
-
GPT-5.5 Achieves Breakthrough in Long-Running Task Reliability
By
–
GPT-5.5 is MUCH more reliable on longer running tasks – for the first time with any model.
— Peter Gostev (@petergostev) 23 avril 2026
As we speak I have a migration running for over 7+ hours – this literally never happened before, the models would maybe run for 30 mins or of you really shout at them for 2-3 hours.… pic.twitter.com/nxdrKZnOzWGPT-5.5 is MUCH more reliable on longer running tasks – for the first time with any model. As we speak I have a migration running for over 7+ hours – this literally never happened before, the models would maybe run for 30 mins or of you really shout at them for 2-3 hours.
-
GPT-5.5 Demonstrates Significantly Enhanced Capabilities Over Previous
By
–
When using GPT-5.5, it is instantly noticeable how much more powerful it is.
— Peter Gostev (@petergostev) 23 avril 2026
In Codex, I gave it a very complex prompt to create London Toy Railway with landmarks and seasons – it did an excellent job in one shot.
In the second half of the video you see GPT-5.4 – it was also… pic.twitter.com/zmAgrAsYilWhen using GPT-5.5, it is instantly noticeable how much more powerful it is. In Codex, I gave it a very complex prompt to create London Toy Railway with landmarks and seasons – it did an excellent job in one shot. In the second half of the video you see GPT-5.4 – it was also
-

ChatGPT analyzes spending patterns to predict future financial situations
By
–
Oh wow, if you feed your bank statements to ChatGPT, it can show you what you'd look like in 10 years time if you maintain your level of spending
-
Balancing latency considerations in AI generation systems
By
–
For now the consideration was latency, it could take a few minutes to generate, but we'll see if it would be worth doing on balance
-
ChatGPT Thinking vs API: How AI Models Process Information
By
–
It was Medium and I'm not 100% sure how thinking works in ChatGPT and whether it was equivalent – I would say it is equivalent to how the API works, so I think ChatGPT thinking would be different.
-

Image Quality Comparison: High Resolution vs Low Resolution AI Outputs
By
–
PSA: High (4k) vs Low (1k) quality: both are not bad, but if you zoom in, the high quality has substantially more detail and no errors, while the low quality very notesable issues with text rendering. ChatGPT itself has another layer of complexity for Thinking, but these are
-
Resolution Impact on AI Image Generation Quality and Output
By
–
Yes resolution could impact the quality of the output too – basically if it generates more pixels it has more opportunity to do a better job, so this isn't just an upscale. But they are also saying it is an experimental feature, so could create random issues too potentially
-
Professional AI Users Seeking to Test Model Token Limits
By
–
As a PROfessional user of AI, what I'm looking for is to go to a website, put in a prompt and hit the limit before it gets to an answer
-

Optimize Image Quality Settings for AI Generation Tools
By
–
One important note – there is a 'quality' setting and this looks like 'low' to me, I don't know how this gets handled in chatgpt, but on the playground you can select the 'high' setting and '4k' resolution, it will be a lot better I promise, I did like a 100 of these.
