OpenAI o1 — our first model trained with reinforcement learning to think hard about problems before answering. Extremely proud of the team! This is a new paradigm with vast opportunity. This is evident quantitatively (eg reasoning metrics are already a step function improved)
OpenAI o1: Reinforcement Learning Breakthrough for Advanced Reasoning
By
–