We used a simple yet powerful approach: We simultaneously generated multiple candidate solutions using GPT-5 and an internal experimental reasoning model, then used our experimental model to intelligently select the optimal solutions for submission. There was no complex strategy
GENERATIVE AI
-

AI Solves All ICPC World Finals Problems Achieving First Place
By
–
Our general-purpose reasoning models solved all 12 problems at the 2025 International Collegiate Programming Contest (ICPC) World Finals, the world’s top university programming competition which was enough for a 1st-place human ranking.
-
OpenAI o1 Models Achieve Top Marks in Math Coding Competitions
By
–
This caps a run of steady progress across math and coding competitions. Just over a year ago we introduced OpenAI o1-preview and OpenAI o1-mini. Since then our general-purpose reasoning models have made steady progress. Today they’re earning top marks in some of the world’s
-

Self-Correcting AI Agents Enable Exponential Task Horizon Gains
By
–
I think the significance of this is under-appreciated: the assumption has often been that AI agents are brittle as one failure in a chain breaks a task But this paper shows smart models are self-correcting & that small gains in accuracy lead to exponential gains in task horizons
-
Gemini 2.5 Deep Think Achieves Gold-Medal Performance at ICPC World Finals
By
–
Very excited to see our Gemini models getting better and better at coding! An advanced version of Gemini 2.5 Deep Think at the 2025 International Collegiate Programming Contest (ICPC) World Finals achieved gold-medal level performance!
-

OpenAI Wins Math Competition Solving All 12 Problems
By
–
Otra competición de matemáticas donde OpenAI participa y logra oro superando en este caso todos los problemas (12 de 12)! No han usado un modelo entrenado específicamente para esto sino un combo de modelos generales. Entre ellos el misterioso modelo avanzado de matemáticas
-
100+ Free Step-by-Step AI Agents and RAG Tutorials
By
–
100+ free step-by-step tutorials with code covering: AI Agents RAG Systems Voice AI Agents MCP AI Agents Multi-agent Teams Autonomous Game Playing Agents P.S: Don't forget to subscribe for FREE to access future tutorials.
-

Alibaba releases 30B agentic LLM outperforming Claude and DeepSeek
By
–
China's Alibaba just dropped an opensource 30B agentic LLM that outperforms Claude 4 Sonnet, DeepSeek v3.1, Kimi k2 on a range of agentic search benchmarks. Only ~3B parameters are activated per token. 100% Opensource.
-

Frontier AI Models Show Scheming Behaviors, Explicit Reasoning Reduces Risk
By
–
In this new research with @apolloaievals
, we found behaviors consistent with scheming in controlled tests across frontier models, including OpenAI o3 and o4-mini, Gemini-2.5-pro, and Claude Opus-4. We can significantly reduce scheming by training models to reason explicitly, -
Frontier Models Show Scheming Behaviors, Mitigation Strategy Tested
By
–
Today we’re releasing research with @apolloaievals
. In controlled tests, we found behaviors consistent with scheming in frontier models—and tested a way to reduce it. While we believe these behaviors aren’t causing serious harm today, this is a future risk we’re preparing