

OpenAI is testing the "robin-high" model on LM Arena. It is the only model on LM Arena to get this math question right besides Gemini 3. Looks like I need to upgrade my questions

By
–


OpenAI is testing the "robin-high" model on LM Arena. It is the only model on LM Arena to get this math question right besides Gemini 3. Looks like I need to upgrade my questions

By
–



BREAKING : OpenAI hints at “garlic”, a rumoured codename of their upcoming model. GPT-5.2 is expected tomorrow

By
–
Google updated Gemini 2.5 Flash and Pro Text-to-Speech (TTS) models with new capabilities. – Emotional style and tone versatility – Context-aware pacing control – Improved multiple-speaker capabilities Both models replaced older versions in AI Studio
By
–
Yutory released Scouts, a team of AI agents that can monitor information on web resources for you.
— 🚨 AI News | TestingCatalog (@testingcatalog) 10 décembre 2025
For example, it can scout for the latest AI news from TestingCatalog and notify you as soon as a new post comes out.
Or any other website 👀 https://t.co/slpU9dOt67 pic.twitter.com/UjfOnKhMbQ
Yutory released Scouts, a team of AI agents that can monitor information on web resources for you. For example, it can scout for the latest AI news from TestingCatalog and notify you as soon as a new post comes out. Or any other website
By
–
Google upgraded Stitch design Agent with Gemini 3 Pro, which is the new default mode now.
— 🚨 AI News | TestingCatalog (@testingcatalog) 10 décembre 2025
Testing time 👀 https://t.co/uW1w4r2fOe pic.twitter.com/oAaZrtvFTe
Google upgraded Stitch design Agent with Gemini 3 Pro, which is the new default mode now. Testing time
By
–
Cursor released a new Debug mode, which can analyse server logs and instrument the code in order to find a problem.
— 🚨 AI News | TestingCatalog (@testingcatalog) 10 décembre 2025
What else? 👀
– Plan Mode now supports Mermaid diagrams
– Multi-Agent Judging
– Pinned Chats https://t.co/tL37cE7a1e pic.twitter.com/ue2szZ2a5f
Cursor released a new Debug mode, which can analyse server logs and instrument the code in order to find a problem. What else? – Plan Mode now supports Mermaid diagrams
– Multi-Agent Judging
– Pinned Chats
By
–
BREAKING 🚨: Google released Scheduled Tasks and Suggestions in Beta on Jules SWE Agent, where it can proactively propose improvements to your repository.
— 🚨 AI News | TestingCatalog (@testingcatalog) 10 décembre 2025
The whole UI got redesigned as well 🤩 https://t.co/ZhJkaTMN9k pic.twitter.com/cto6tHnouq
BREAKING : Google released Scheduled Tasks and Suggestions in Beta on Jules SWE Agent, where it can proactively propose improvements to your repository. The whole UI got redesigned as well
By
–
"For example, capabilities assessed through capture-the-flag (CTF) challenges have improved from 27% on GPT‑5 in August 2025 to 76% on GPT‑5.1-Codex-Max in November 2025."
By
–
Meta is working on a new proprietary frontier AI model called “Avocado” according to CNBC. “We were expecting the model to be released before the end of this year, but that the plan now is for that to happen in the first quarter of 2026.” Meta will be back

By
–
BREAKING : Notion might be testing GPT-5.2 internally under the "olive-oil-cake" codename. The codename is abstracted by Notion. Soon or not soon?