Ok I was wrong – 5.4 Mini and 5.4 Nano are out. Mini is 1/3 of 5.4 and nano is 1/10 the price
@petergostev
-

Grok 4.2 Multi-Agent Token Efficiency Benchmark Update
By
–
That reasoning basically doesn't help at all, the only benchmark I can think of that shows this. Maybe I need to update this chart as Grok 4.2 multi agent ate up way too many tokens
-

Comparing AI Model Performance: Sonnet vs Opus Benchmarks
By
–
The label didn't fit in, it is slightly below Sonnet and Opus 4.5 – but I wouldn't read it too precisely, they are all about same ballpark, see here: https://
petergpt.github.io/bullshit-bench
mark/viewer/index.v2.html
… -
Who Is Jason and Why Is He Deleting Production Database
By
–
Ok I'll bite, who is Jason and why is he deleting my prod db?
-
Opus Audit Assessment of Leading AI Models Compared
By
–
I was using Opus via Cursor, did an audit with Gemini 3.1 Pro, Opus 4.6 and GPT-5.4. Then I asked Opus to give assessment of the audit quality (anonymously). And I think it 100% nailed the current state of the models: Gemini 3.1 Pro: The weakest. Looked at the screen. Found the
-

Codex Issues and Account Reset Considerations for Users
By
–
Any Codex niche issues warranting a reset on the horizon @thsottiaux
? My Pro account is suffering -

Grok 4.2 Achieves Major Rankings Improvement on BullshitBench
By
–
BullshitBench v2 Update: Grok 4.2 – massive jump in the rankings – 4.1 was ranked 54th and 72nd (out of 84) and now it took 13-16th spots. https://t.co/EdwOBStKa4 pic.twitter.com/6T6Tr0KsDr
— Peter Gostev (@petergostev) 12 mars 2026BullshitBench v2 Update: Grok 4.2 – massive jump in the rankings – 4.1 was ranked 54th and 72nd (out of 84) and now it took 13-16th spots.
-
AI Agents Need General Computer Use Beyond Browser Limitations
By
–
Yeah I think you are right, I wish there was a general computer use capability and not browser limited
-
Playwright Interactive Skill Now Working Well
By
–
Have you tried the playwright interactive skill? Seems to be working quite well
-
Codex Model Efficiency Falls Short of OpenAI Expectations
By
–
Interesting that it wasn't more efficient than 5.3 codex, like OpenAI said it should be
