Good point. Much reduced hallucination is so underrated. I don't understand why so few people talk about it.
LLMS
-
GPT-5 Release Meets Expectations But Shows Disappointing Aspects
By
–
My take: it meets my expectations exactly. Since there was a lot of discussion about GPT-5 beforehand, my expectations are pretty much in line with how it was finally released. At the same time, there are things that disappoint me and are underwhelming (instruction following,
-
GPT-5 Performance Assessment: Expectations vs Reality
By
–
GPT-5: did it exceed, meet or fall short of your expectations? Please feel free to provide arguments 🙂
-

GPT-5 vs GPT-5 mini vs GPT-5 nano: Quick pick guide
By
–
GPT-5 vs GPT-5 mini vs GPT-5 nano
— Louis-François Bouchard 🎥🤖 (@Whats_AI) 8 août 2025
Don’t default to GPT-5. Here’s the quick pick guide 👇
What they all share:
• 400k context
• images + text input, text output only
• tool access
• “Minimal reasoning” toggle = lower cost
Use GPT-5 when…
• You need max reasoning/agent… pic.twitter.com/5Lm5wjLStzGPT-5 vs GPT-5 mini vs GPT-5 nano Don’t default to GPT-5. Here’s the quick pick guide What they all share:
• 400k context
• images + text input, text output only
• tool access
• “Minimal reasoning” toggle = lower cost Use GPT-5 when…
• You need max reasoning/agent -

CoAct-1: Computer-Using Agents with Coding Actions
By
–
CoAct-1 Computer-using Agents with Coding as Actions
-

R-Zero: Self-Evolving Reasoning LLM Without Initial Data
By
–
R-Zero Self-Evolving Reasoning LLM from Zero Data
-

Generalization of SFT: Reinforcement Learning with Reward Rectification
By
–
On the Generalization of SFT A Reinforcement Learning Perspective with Reward Rectification
-

GPT-5 Launch, TRISO Fuel, and DeepMind Genie 3 Highlight Week
By
–
DOE taps Standard Nuclear for TRISO.
General Matter opens KY uranium enrichment facility.
GPT-5 is here & humans are still useful.
Google DeepMind stuns with Genie 3.
Eleven Labs drops ElevenMusic.
Lithium shows promise in Alzheimer's prevention. What a week for the optimists. -
Testing o3 Performance Against Itself for Comparison
By
–
o3 was already outstanding. I need to do more testing to see if it is better than o3 in my experience.
-
GPT-5 Shows No Improvement Over GPT-4o in Instruction Following
By
–
From my initial testings GPT-5 is as bad as GPT-4o on instruction following. This is a real bummer. No improvement at all – but I need more testing to validate this.