BullshitBench update: Gemini 3.5 Flash did pretty badly – below a bit Gemma 4 even (31b high is 80th)
@petergostev
-
EU AI regulation critique and preemptive harm concerns
By
–
I don't believe pre agreeing anything works – eg EU did AI regulation which they started drafting before ChatGPT, as a result it is pretty much harmful without any of the benefits. Preparing some framework before we know what the world looks like could be harmful too. For
-

GPT Image 2 sees 50% usage growth and 1.5bn weekly images
By
–
GPT Image 2: two weeks after launching the model, usage is up 50% and 1.5bn images generated every week in ChatGPT only.
-
Coding with o3 model to avoid AI slowdown
By
–
If you feel like AI is slowing down, try coding with o3 – best model only a year ago
-
Using AI agents for migration task documentation
By
–
The goal, really? I have quite a big migration task and I was working in one thread to work out how to set it up well – e.g. the target spec, testing approach, etc. then got codex to write out docs and the goal prompt for the other agent. Then I'm periodically checking in with
-

AI goal-setting feature in Codex app
By
–
Did you know that /goal already exists in the codex app, there's no UI around it, but if you literally just start a message with /goal and write your goal then you can get it in the app, not just the CLI
-

The Impact of Prompt Caching on LLM Agentic Workflows and Costs
By
–
Prompt caching didn't even exist until <2 years ago Google: June 2024
Anthropic: August 2024
OpenAI: October 2024 Now for agentic workflows, 95% of tokens gets cached, without it the costs would be completely insane -
Challenges in Enterprise AI Adoption and Employee Usage
By
–
I know CEO-level top down tokenmaxxing is considered dumb, but what are they to do when they've exhausted 'the good way' – you got everyone licenses, did AI workshops, appointed AI champions, did AI hackathons – then you check back and adoption is 25%, of which 80% is tab
-
Voice-first shopping for dino outfits
By
–
I've always wondered why we don't see any voice-first interfaces, so I tried to build one with OpenAI's new realtime voice api – you can now shop for your dino outfits with just your voice. pic.twitter.com/gKFekQ84RN
— Peter Gostev (@petergostev) 10 mai 2026I've always wondered why we don't see any voice-first interfaces, so I tried to build one with OpenAI's new realtime voice api – you can now shop for your dino outfits with just your voice.
-
Even image models fail as world models; video is harder
By
–
Even the image models are not good enough world models yet, gpt-image-2 and nano banana make pretty obvious world model like mistakes. is way harder, so no hope of that anytime soon. Maybe with 2-3 OOMs it gets good enough, but that doesn't feel like a reasonable thing to
