not sure what you are saying here. i don’t see data showing hallucinations have been solved, as you are asserting.
RESEARCH
-
Gemini-3-Pro Hallucinations Persist Despite Claims Study
By
–
1. you have given me zero studies with whatever models you think are relevant supporting your claim that X LLMs no longer hallucinate.
2. Gemini-3-pro hallucinates in https://
arxiv.org/pdf/2603.21687 just a couple weeks ago
3. Turning this around into a question about my own use of models -

OpenAI Safety Concerns Revealed by New Yorker Investigation
By
–
This news comes hours after @NewYorker published its investigation detailing the various ways AI experts warn OpenAI hasn't been taking AI safety seriously enough.
-
Recent Frontier Model Data Availability Questions
By
–
its an older paper. do you have more recent data on frontier models you aren’t sharing?
-
Benchmarks and hallucinations in reasoning models clarified
By
–
that’s one particular benchmark that came up; not a universal claim by me (or anyone else AFAIK). compare “Gary mentioned a particular benchmark” that didn’t include reasoning models (true) with “no studies have indicated hallucinations within reasoning models” (false). you
-
Stabilizing Video from Running Animals: New Petpin Pipeline Breakthrough
By
–
Stabilizing video from a camera on a running animal turns out to be brutally hard.
— Ark Baltser (@arkslife) 6 avril 2026
Traditional stabilization breaks pretty quickly.
We're starting to crack it.
Before → After from our latest Petpin pipeline. pic.twitter.com/L4NWftjfxCStabilizing video from a camera on a running animal turns out to be brutally hard. Traditional stabilization breaks pretty quickly. We're starting to crack it. Before → After from our latest Petpin pipeline.
→ View original post on X — @scobleizer, 2026-04-06 20:17 UTC
-
LRMs expose blind spot in reasoning capabilities assessment
By
–
One funny thing about the recent rise of LRMs is that the people who were adamant that base LLMs from 2023-2024 could already reason completely missed it, as they didn't know what to look for. You can't notice something you don't expect.
-
LLMs Auto-Regressive Nature: Beyond AI Hype
By
–
Guessing you're referring to the story and not my tweet, but totally agree that everyone should remember LLMs are by design auto-regressive. (Shocking so many still don't.) AI companies should stop over-hyping LLMs and start explaining how they actually work and where they're
-
Base LLMs Fail at Math, LRMs Make Progress
By
–
Paper below tested a variety of base LLMs (no TTA) on generalization-focus math problems and found that they can't reason and can't do math. All true… but the fact that base LLMs have zero fluid intelligence, while extremely controversial back in 2024, is now well established. An interesting experiment here would have been to try current LRMs on the same problems and measure the delta. I bet latest LRMs can solve most of these problems. arxiv.org/abs/2604.01988 [Translated from EN to English]
-
Can AI Build a Simple Single-Purpose App Correctly?
By
–
I’m talking about a very simple app that does just one thing and does it correctly. Something that even *I* could develop if someone told me what it is. If AI can’t do it then IMHO fails a very very straightforward test of intelligence.