Cool figure of #Gemini3 with #DeepThink on ARC-AGI-2. The big jump looks quite like the results on #IMOBench that we obtained previously with Gemini 2.5
@lmthang
-

LLM Achieves #1 Ranking Across All Arena Leaderboards
By
–
And #1 across all LLM @arena leaderboards! I guess we were kind enough to not be on the leaderboard for a day π
-

IMO-Bench Development: Gemini Math Capabilities Evolution
By
–
A bit of a history of IMO-Bench and our IMO efforts:
a. We started building IMO-Bench around early 2024, which was the precursor of ProofBench (basic). b. IMO-Bench was first mentioned in the Gemini 1.5 paper around May 2024. At that time, Gemini Math-specialized 1.5 Pro scored -
AI Research Breakthrough with International Mathematics Experts
By
–
Finally, this work won't be possible with the co-authors: @DawsenHwang
, @HoangT1215
, @g01na2
, @junsukim
, @gjb_ai
, @jon_lee0
, @Swarooprm7
, @hmichalewski
, @XingyouSong
, @thtrieu_
, @quocleix
, @jj_at_brown
, our IMO experts (who account for a total of 10 gold and 5 silver IMO medals), -

GradingBench releases 1000 human-graded IMO proof data points
By
–
Last but not least, #GradingBench has 1000 human grading data points from our IMO effort on the advanced IMO-ProofBench. To ensure a robust evaluation, the dataset has been balanced across 4 simplified grading categories (Correct, Almost, Partial, Incorrect). We released all
-

IMO-Bench: 400 olympiad problems for AI evaluation
By
–
IMO-Bench also includes #AnswerBench, a set of 400 problems with verifiable answers carefully chosen from past Olympiad competitions. The problems span across four IMO categories (Algebra, Combinatorics, Geometry, and Number Theory), were altered by experts to avoid memorization,
-
Gemini Deep Think achieves strong performance on FrontierMath benchmark
By
–
Hopefully, this thread is a good teaser to the strengths of Gemini Deep Think (IMO-gold) given that last month, Gemini Deep Think (IMO-lite) topped FrontierMath π
-

Grok 4 Overfitting Analysis ProofBench Advanced Leaderboard
By
–
For those of you who love leaderboards, here is one for the advanced #ProofBench π When breaking down performances into different subsets, we noticed potential overfitting in certain models and approaches. For example, Grok 4 (heavy) scores 76.2% on USAMO 2025 but only 11.1% on
-

IMO Medalists Grade AI Homework with Human Verification
By
–
We do have teachers (IMO medalists) to grade the homeworks too π See the paper https://
arxiv.org/abs/2511.01846 where we recommend to augment with human verifications. -

ProofAutoGrader: Automatic IMO Proof Evaluation Using Gemini
By
–
While human expert evaluation remains the gold standard for mathematical proofs, its cost and time intensity limit scalable research. To address this, we built #ProofAutoGrader, an automatic grader for IMO-ProofBench. The autograder leverages Gemini 2.5 Pro, providing it with a
