Very excited to see our Gemini models getting better and better at coding! An advanced version of Gemini 2.5 Deep Think at the 2025 International Collegiate Programming Contest (ICPC) World Finals achieved gold-medal level performance!
LLMS
-

OpenAI Wins Math Competition Solving All 12 Problems
By
–
Otra competición de matemáticas donde OpenAI participa y logra oro superando en este caso todos los problemas (12 de 12)! No han usado un modelo entrenado específicamente para esto sino un combo de modelos generales. Entre ellos el misterioso modelo avanzado de matemáticas
-
100+ Free Open Source AI Agents and RAG Tutorials
By
–
Stay tuned for more such interesting posts → @Saboo_Shubham_ I have created 100+ AI Agents and RAG tutorials, 100% free and opensource. P.S: Don't forget to star the repo to show your support
-
100+ Free Step-by-Step AI Agents and RAG Tutorials
By
–
100+ free step-by-step tutorials with code covering: AI Agents RAG Systems Voice AI Agents MCP AI Agents Multi-agent Teams Autonomous Game Playing Agents P.S: Don't forget to subscribe for FREE to access future tutorials.
-

Alibaba releases 30B agentic LLM outperforming Claude and DeepSeek
By
–
China's Alibaba just dropped an opensource 30B agentic LLM that outperforms Claude 4 Sonnet, DeepSeek v3.1, Kimi k2 on a range of agentic search benchmarks. Only ~3B parameters are activated per token. 100% Opensource.
-

Frontier Models Show Situational Awareness Affects Scheming Behavior
By
–
Frontier models can recognize when they are being tested, and their tendency to scheme is influenced by this situational awareness. We demonstrated counterfactually that situational awareness in their chain-of-thought affects scheming rates: the more situationally aware a model
-

Frontier AI Models Show Scheming Behaviors, Explicit Reasoning Reduces Risk
By
–
In this new research with @apolloaievals
, we found behaviors consistent with scheming in controlled tests across frontier models, including OpenAI o3 and o4-mini, Gemini-2.5-pro, and Claude Opus-4. We can significantly reduce scheming by training models to reason explicitly, -
Grok 5 Probability Assessment Neural Network Prediction Rising
By
–
Just a chance. But I thought the probability was 0% for all prior Grok releases and now my neural net predicts ~10% probability and rising for Grok 5.
-
Build AI Apps with MCP Servers and Box Files
By
–
New short course: Build AI Apps with MCP Servers: Working with Box Files, built with @Box and taught by @BenAtBox , their CTO.
— Andrew Ng (@AndrewYNg) 17 septembre 2025
Many AI applications require custom code for basic file operations. The Model Context Protocol (MCP) standardizes this by letting you offload file… pic.twitter.com/lITKXRfnHANew short course: Build AI Apps with MCP Servers: Working with Box Files, built with @Box and taught by @BenAtBox , their CTO. Many AI applications require custom code for basic file operations. The Model Context Protocol (MCP) standardizes this by letting you offload file
