first, the problem quantified. the researchers created a metric called RFCS (Ratio of First Correct Step) that tracks where in a chain of thought the correct answer first appears. on MATH-500, across every model tested, the right answer shows up well before the end in over half
LLMS
-

Overthinking in AI: A Sampling Issue
By
–
reasoning models already know when they've solved the problem. we just don't let them stop. new paper from Beihang University and ByteDance shows that the overthinking problem in models like DeepSeek-R1 and Qwen3 isn't a training failure. it's a sampling failure. the fix cuts
-
GPT-5.4 outperforms GPT-5.2 on code and knowledge benchmarks
By
–
that’s not true, evals comparing 5.2 vs 5.4 (thinking only) Coding (SWE-Bench Pro) > GPT-5.2: 55.6%
> GPT-5.4: 57.7%
→ ~+2.1 pts improvement in solving real-world repo bug-fix tasks. Knowledge-work benchmark (GDPval) > GPT-5.2: ~71% win/tie vs professionals
> GPT-5.4: 83% -

Google/Harvard/CMU neuro-symbolic agent uses Gemini to discover proofs
By
–
Google just solved a theoretical physics problem using Gemini! Google, Harvard, and CMU built a neuro-symbolic system using the Gemini Deep Think model and a tree-search framework to autonomously discover complex mathematical proofs. The agent functions like a digital
-

Comparison between Claude Code and Codex App
By
–
Claude Code vs Codex App [Translated from EN to English]
→ View original post on X — @arrakis_ai, 2026-03-06 07:44 UTC
-
Recursive Behaviors Emerge in Frontier AI Models Without RLM Training
By
–
What I liked about thsi paper was that these recursive behaviors emerge without explicit RLM training. This is because frontier models like GPT-5, Qwen3-Coder, etc., already have enough computational intuition to grep, partition, and spawn sub-calls effectively. So they just
-

NEO-unify: Building Native Multimodal Unified Models End to End
By
–
NEO-unify: Building Native Multimodal Unified Models End to End Blog: https://
huggingface.co/blog/sensenova
/neo-unify
… -

OpenAI announces GPT-5.4 with 1M token context and extreme reasoning
By
–
And same in bullet points, thanks to @blevlabs
's AI agent he and I built together. Here's what OpenAI announced today (March 5, 2026):
• GPT-5.4 launched — new frontier model with 1 million token context window, "extreme" reasoning mode, and the ability to interrupt the model -

OpenAI announces GPT-5.4 with 1M token context and new features
By
–
Here you go, thanks to https://
levangielabs.com Here's what OpenAI announced today (March 5, 2026):
• GPT-5.4 launched — new frontier model with 1 million token context window, "extreme" reasoning mode, and the ability to interrupt the model mid-response to redirect it -
GPT 5.4 release impact on developers and AI agents video
By
–
What does @OpenAI GPT 5.4 (just released this morning) mean for developers?
— Robert Scoble (@Scobleizer) 6 mars 2026
You talked about it. My AI agents read it. Wrote a report. Sent it over to @NotebookLM. Which made this video. Using the new cinematic quality released yesterday.
This was all built with your posts from… pic.twitter.com/4SHT2IGd1oWhat does @OpenAI GPT 5.4 (just released this morning) mean for developers? You talked about it. My AI agents read it. Wrote a report. Sent it over to @NotebookLM
. Which made this video. Using the new cinematic quality released yesterday. This was all built with your posts from