Thanks! In longer chats I'm convinced models respond to messages 1 in the past, maybe due to timeout/revert earlier in the conversation. I have been sending them message codes they have to echo back, and sometimes comes back 1 delayed, response content is also off-by-one. Could
AGI
-
AI system outperforms humans in R&D, tackling humanity’s biggest problems.
By
–
Autonomous research agents have the potential to solve some of humanity’s biggest problems.
— AI Breakfast (@AiBreakfast) 19 novembre 2025
This is the first AI system to outperform humans in R&D – very cool! https://t.co/UTeBBkfthKAutonomous research agents have the potential to solve some of humanity’s biggest problems. This is the first AI system to outperform humans in R&D – very cool!
-
LLMs Creating Superior RL Environments for Advanced Learning
By
–
LLMs will be creating RL environments far superior to any software tools humans have ever created.
— Nando de Freitas (@NandoDF) 19 novembre 2025
LLMs will use their created tools for learning , and for greater sensing and control.
It’s coming, as this leap by Gemini 3 shows. https://t.co/cWXyqO1H8lLLMs will be creating RL environments far superior to any software tools humans have ever created. LLMs will use their created tools for learning , and for greater sensing and control. It’s coming, as this leap by Gemini 3 shows.
-
Scaling Reinforcement Learning Critical Path to AGI
By
–
This didn't age well Kidding obviously, keep up the good work. Scaling RL will be critical towards AGI.
-
Thinking Machines: The Incredible People Behind AI
By
–
thinking machines….the people are incredible
-

Google Gemini 3.0 Benchmarks Rival Top AI Models
By
–
Google sort Gemini 3.0, 24h après Elon Musk ! Les benchmarks de Gemini 3 Deep Think sont absolument fous. – HLE (raisonnement & connaissances, sans outils) : 41 %
– GPQA Diamond (connaissances scientifiques, sans outils) : 93,8 %
– ARC-AGI-2 (puzzles de raisonnement visuel, -

Gemini 3 DeepThink Shows Major ARC-AGI Performance Leap
By
–
Cool figure of #Gemini3 with #DeepThink on ARC-AGI-2. The big jump looks quite like the results on #IMOBench that we obtained previously with Gemini 2.5
-

Gemini 3 Deep Think Achieves 45.1% on AGI Benchmark
By
–
WTF.. Gemini 3 Deep Think hit 41 percent on HLE and reached 45.1 percent on the ARG-AGI-2 benchmark.
