I have an ML professor asserting that hallucinations have been solved in recent frontier reasoning models. Please supply any data you know that supports or rebuts this claim.
→ View original post on X — @garymarcus, 2026-04-06 20:39 UTC
By
–
I have an ML professor asserting that hallucinations have been solved in recent frontier reasoning models. Please supply any data you know that supports or rebuts this claim.
→ View original post on X — @garymarcus, 2026-04-06 20:39 UTC
By
–
not sure what you are saying here. i don’t see data showing hallucinations have been solved, as you are asserting.
By
–
1. you have given me zero studies with whatever models you think are relevant supporting your claim that X LLMs no longer hallucinate.
2. Gemini-3-pro hallucinates in https://
arxiv.org/pdf/2603.21687 just a couple weeks ago
3. Turning this around into a question about my own use of models
By
–
its an older paper. do you have more recent data on frontier models you aren’t sharing?
By
–
that’s one particular benchmark that came up; not a universal claim by me (or anyone else AFAIK). compare “Gary mentioned a particular benchmark” that didn’t include reasoning models (true) with “no studies have indicated hallucinations within reasoning models” (false). you
By
–
Yeah, and soon Elon says it will be able to read lists, create scripts, and create TV shows and other stuff from the lists. I expect that by the end of the year.
By
–
One funny thing about the recent rise of LRMs is that the people who were adamant that base LLMs from 2023-2024 could already reason completely missed it, as they didn't know what to look for. You can't notice something you don't expect.
By
–
Guessing you're referring to the story and not my tweet, but totally agree that everyone should remember LLMs are by design auto-regressive. (Shocking so many still don't.) AI companies should stop over-hyping LLMs and start explaining how they actually work and where they're
By
–
Codex for open source update! Some of the main projects we’ve supported: – Linux – React – Node.js – Rust – Python / CPython – Kubernetes – Flutter – Electron – Ollama – Dify – Transformers – LangChain – yt-dlp – OpenCV – Home Assistant – Storybook – Astro – vLLM – SGLang – TiDB – Airflow – Superset – Doris – APISIX – Pulsar – RocketMQ – DataFusion – Arrow – SeaTunnel – ShardingSphere – DolphinScheduler – Apache ECharts – Ant Design – Tailwind CSS – Material UI – Svelte – Vue – Solid – Qwik – Hono – tRPC – Prisma – TypeORM – NextAuth – TanStack Query – TanStack Table – Rspack – Biome – Ruff – Docusaurus – Zed – Tauri – Deno – Turso – SQLite Browser – Difftastic – core-js – nvm – pnpm – zoxide – fd – bat – k9s – Harbor – containerd – etcd – Envoy – gRPC – Jaeger – Trivy – Vaultwarden – Firecracker – Gitea – Homebrew Core – Nixpkgs – Julia – Ruby – Rails – Symfony – scikit-learn – PaddlePaddle – Diffusers – Llama Factory – Unsloth – DeepSpeed – DSPy – RAGFlow – MinerU – ComfyUI – Jan – Khoj – LibreChat – Chainlit – LobeHub – GPT Engineer – Browser Use – OpenClaw – Zeroclaw – Plane – Appwrite – Directus – Strapi – Calcom – Metabase – PostHog – Budibase – Documenso – Chatwoot – Logto – Zammad – Discourse – BTCPay Server – Bagisto – ABP – Vitess – TiKV – RisingWave – Milvus – Kong – SigNoz – Manticore Search – JuiceFS – SeaweedFS – Reth – Aptos Core – Ripple – Wormhole – Go Ethereum – Bevy – Godot – PixiJS – Excalidraw – JSON Crack – RSSHub – SiYuan – Logseq – Refined GitHub – Better Auth – OpenPilot – FFmpeg – XBMC / Kodi – MeshCentral – ZMap – Yacy – Zplug – ZipArchive – ZIO – Godoxy – rengine – pyod – node-telegram-bot-api – Avante.nvim – Repomix – Zfile – Vant – Vant Weapp – Yamada UI – WordPress Playground – Gutenberg And a bunch of number of Apache, AI, database, frontend, security, CLI, infra, and devtool projects!
→ View original post on X — @romainhuet, 2026-04-06 20:17 UTC
By
–
Paper below tested a variety of base LLMs (no TTA) on generalization-focus math problems and found that they can't reason and can't do math. All true… but the fact that base LLMs have zero fluid intelligence, while extremely controversial back in 2024, is now well established. An interesting experiment here would have been to try current LRMs on the same problems and measure the delta. I bet latest LRMs can solve most of these problems. arxiv.org/abs/2604.01988 [Translated from EN to English]