IMO-ProofBench is our key focus designed to evaluate the ability of AI models in constructing rigorous and valid mathematical arguments. With 60 proof-based problems, the benchmark is divided into two subsets: a basic set covering pre-IMO to IMO-Medium difficulty levels, and an
@lmthang
-

IMO-Bench: New AI Model Evaluation Framework for Mathematics
By
–
IMO-Bench consists of three benchmarks that judge models on diverse capabilities: IMO-AnswerBench, a large-scale test on getting the right answer; IMO-ProofBench, a next-level evaluation for proof writing; and IMO-GradingBench, a new benchmarkto enable further progress in
-
DeepThink IMO Lite Now Available Public Gemini Ultra
By
–
Correction: it's a public model access through Gemini Ultra subscription (not API). I guess the main point is DeepThink IMO lite is available for the public to assess 🙂
-
Gemini Ultra Subscription Now Offers Public Model Access
By
–
Yea, sorry publicly available model access through Gemini Ultra subscription
-

DeepThink IMO Gold Model Achieves State-of-the-Art FrontierMath Results
By
–
Very cool to see this new state-of-the-art result on FrontierMath achieved by the #DeepThink IMO Gold model that we built 3 months (a long time) ago 🙂 The fact that is evaluated externally using a publicly available API is a strong testament that what we built generalizes beyond
-

Gemini DeepThink Wins Gold at ICPC2025 Programming Contest
By
–
Yet another important milestone for Gemini #DeepThink in achieving goldat #ICPC2025, the world's most prestigious college-level programming contest. Big congrats to everyone in this great achievement! Stay tuned for more to come from DeepThink 😉
-
Google Gemini DeepThink Model Receives External Validation
By
–
Glad to hear external validations like this for the Gemini #DeepThink IMO model we shipped last month 🙂 Stay tuned for future updates!
-

NewTuring nurtures AI talents benchmark achievement Asia-Pacific
By
–
Congrats @longphan3110 on yet another cool benchmark (after HLE). Long grew up from @newturing (formerly VietAI) working with me & @thtrieu_ on machine translation. Now return back to https://t.co/fRvhGyWatj to help nurture young talents from Asia-Pacific (<12hrs left to apply!) https://t.co/oMgK95hRXG
— Thang Luong (@lmthang) 13 août 2025Congrats @longphan3110 on yet another cool benchmark (after HLE). Long grew up from @newturing (formerly VietAI) working with me & @thtrieu_ on machine translation. Now return back to https://
gstar.newturing.ai to help nurture young talents from Asia-Pacific (<12hrs left to apply!) -
Stanford NLP Group Alumnus Highlights Research Excellence
By
–
Thanks! Proud to be an alumnus of the Stanford NLP group 🙂
-
Nurturing Young AI Talents in Southeast Asia
By
–
Also I'm doing this while I'm on a two-week vacation 🙂 The hope is to find and nurture young talents in Southeast Asia and beyond. And hopefully the community will carry on the work further.
