25,000 daily tokens per human alive today – this is about the amount of tokens we are processing today globally
@petergostev
-
LLM Coding Benchmarks Questioned Real World Performance
By
–
These kind of claims never pass the sniff test. Benchmarks can be cheated, but if it worked 0-11% of the time on real tasks (which are not part of benchmarks) nobody would ever use LLMs for coding. https://t.co/zp1qpQjf3P
— Peter Gostev (@petergostev) 19 mars 2026These kind of claims never pass the sniff test. Benchmarks can be cheated, but if it worked 0-11% of the time on real tasks (which are not part of benchmarks) nobody would ever use LLMs for coding.
-
AI Capability Emergence and Future Impact Prediction Challenges
By
–
The reason why it's so difficult to predict the impact of AI is because futures are radically different if certain capabilities emerge fully & quickly or not at all / remain in the 'hack' phase, for example: – Computer use – Taste / judgement – Training or hard to validate
-
Managing AI Agent Threads Requires Intense Cognitive Work
By
–
There's worry that people will stop using their brains with LLMs, but managing several AI agent threads in parallel has been some of the most cognitively intensive work I've done in years
-
GPT-5.4-mini matches fine-tuned GPT-4.1-mini performance
By
–
So far in my tests gpt-5.4-mini matches the performance of my fine tuned gpt-4.1-mini, I'll keep trying to improve it, and it does look like it is better but not a slam dunk yet
-
Bullshit Benchmark: AI Evaluation Dataset and Viewer Tool
By
–
Github: https://
github.com/petergpt/bulls
hit-benchmark
… Data viewer: https://
petergpt.github.io/bullshit-bench
mark/viewer/index.v2.html
… -

Reasoning Approaches Show Limited Effectiveness in AI Systems
By
–
Another confirmation that reasoning doesn't seem to help
-

GPT-5.4 Mini and Nano Score Low on BullshitBench
By
–
BullshitBench update: The new GPT-5.4 mini and nano models score quite low. This screenshot shows OpenAI models only, on the full list would put GPT-5.4-mini around 40th place and Nano is around 70th place. Again thinking didn't help much at all. https://t.co/EdwOBStKa4 pic.twitter.com/mxBGoFUzio
— Peter Gostev (@petergostev) 17 mars 2026BullshitBench update: The new GPT-5.4 mini and nano models score quite low. This screenshot shows OpenAI models only, on the full list would put GPT-5.4-mini around 40th place and Nano is around 70th place. Again thinking didn't help much at all.
-
High Demand Pricing Strategy: OpenAI and Google Market Dynamics
By
–
I think we are at a point where demand is so high that lowering /holding prices just doesn't make sense – same for Google too. If prices are too low, they wouldn't even be able to serve it