Frontier models can’t see, and if you think they can, you’ve probably been fooled by benchmarks that can totally be gamed. In the very short essay linked below I discuss a stunning new finding from Stanford that shows just how serious the problem is. And why this means a lot
RESEARCH
-

Top AI Papers of the Week: March 23-29
By
–
The Top AI Papers of the Week (March 23 – 29) – Claudini
– MemCollab
– ARC-AGI-3
– Composer 2
– Hyperagents
– Attention Residuals
– Agentic AI and the Next Intelligence Explosion Read on for more: -

ChatGPT 26 Times More Likely Give Dangerous Responses Study
By
–
People on this site regularly give me shit, and almost always turn out to be wrong. Like when I said LLMs might well contribute to delusions, and people doubted me. New study shows that ChatGPT was 26 times more likely than a control to give dangerous responses to people
-

Build A Reasoning Model book chapters now available in early access
By
–
It’s done. All chapters of Build A Reasoning Model (From Scratch) are now available in early access. The book is currently in production and should be out in the next months, including full-color print and syntax highlighting. There’s also a preorder up on Amazon.
-
HexRunner Achieves Stable 30 MPH Locomotion Through Speed Design
By
–
Designing for Speed: How HexRunner Achieved Stable 30 MPH Locomotion
— Ronald van Loon (@Ronald_vanLoon) 29 mars 2026
by @lukas_m_ziegler#EmergingTech #Engineering #ArtificialIntelligence #Innovation #Technology pic.twitter.com/W2jAL0BpmWDesigning for Speed: How HexRunner Achieved Stable 30 MPH Locomotion
by @lukas_m_ziegler #EmergingTech #Engineering #ArtificialIntelligence #Innovation #Technology -
Tokens per Watt: The New AI Productivity Metric
By
–
Which is why ‘tokens per watt’ becomes the ultimate measure of productivity, rather than human labour hours. How many units of Intelligence can a watt of electricity produce?
-
AI-Generated Research Published in Nature: Peer Review Implications
By
–
AI Scientist published in Nature is a big deal. I'm curious how the review process handled the fact that the research was AI-generated, that's a fascinating meta question.
-

Crediting Alec for LLM pretraining oversimplifies research history
By
–
Alec is a once in a generational researcher, but saying that he invented pretraining is not only a bit of stretch, but it's also a disrespect to other people's work. Flowers ☾ (@flowersslop) Every LLM from any lab today traces back to this guy, who was the only person at OpenAI pushing for pretraining transformer language models. He built GPT-1. After that did others see the potential. He invented it, and almost none of the so called AI experts even know his name. — https://nitter.net/flowersslop/status/2037892926785634720#m
→ View original post on X — @jeremyphoward, 2026-03-29 12:21 UTC
-
RealityScan Photogrammetry Innovation in Radiance Fields
By
–
Glad you’re pushing the team! @TimSweeneyEpic I hope you will too and ask the RealityScan team to reconsider – RS is so good for photogrammetry and y’all could be at the heart of the radiance field ecosystem too
-
RL Post-training Efficiency: 4x Fewer Rollouts for Coding
By
–
4x fewer rollout turns for competitive accuracy is a big deal for making RL post-training practical. Curious to see if this generalizes beyond coding tasks.