Last but not least, #GradingBench has 1000 human grading data points from our IMO effort on the advanced IMO-ProofBench. To ensure a robust evaluation, the dataset has been balanced across 4 simplified grading categories (Correct, Almost, Partial, Incorrect). We released all
LLMS
-
Gemini Deep Think achieves strong performance on FrontierMath benchmark
By
–
Hopefully, this thread is a good teaser to the strengths of Gemini Deep Think (IMO-gold) given that last month, Gemini Deep Think (IMO-lite) topped FrontierMath 😉
-

Grok 4 Overfitting Analysis ProofBench Advanced Leaderboard
By
–
For those of you who love leaderboards, here is one for the advanced #ProofBench 🙂 When breaking down performances into different subsets, we noticed potential overfitting in certain models and approaches. For example, Grok 4 (heavy) scores 76.2% on USAMO 2025 but only 11.1% on
-

Appreciation for Grok’s Direct Correction Approach
By
–
I appreciate how Grok doesn’t sugar coat corrections.
-

MiniMax-M2 Open Source Model Available on Poe
By
–
MiniMax-M2 is now available on Poe! This open source model has a 200k token context window, has 230b parameters with an MoE architecture, and excels at coding and agent workflows. (1/2)
-
Minimax-M2 AI Model Now Available on Poe Platform
By
–
You can try it at https://
poe.com/Minimax-M2, across the Poe apps on all platforms, and via the Poe API. (2/2) -

IMO-ProofBench: Evaluating AI Mathematical Reasoning Capabilities
By
–
IMO-ProofBench is our key focus designed to evaluate the ability of AI models in constructing rigorous and valid mathematical arguments. With 60 proof-based problems, the benchmark is divided into two subsets: a basic set covering pre-IMO to IMO-Medium difficulty levels, and an
-
Waterfall Development with LLMs: Inherent Challenges
By
–
Waterfall based development with LLMs have all the challenges with waterfall based dev?
-

IMO-Bench: New AI Model Evaluation Framework for Mathematics
By
–
IMO-Bench consists of three benchmarks that judge models on diverse capabilities: IMO-AnswerBench, a large-scale test on getting the right answer; IMO-ProofBench, a next-level evaluation for proof writing; and IMO-GradingBench, a new benchmarkto enable further progress in
-
LangSmith for Improving Agent Quality at Scale
By
–
Why we built LangSmith for improving agent quality
— LangChain (@LangChain) 4 novembre 2025
As more agents move into production, teams need to move beyond vibe-checking and bring rigor to how they understand agent behavior at scale.
In this video, the LangSmith engineering team and Harrison (@hwchase17) sit down to… pic.twitter.com/q7jrXfWJo1Why we built LangSmith for improving agent quality As more agents move into production, teams need to move beyond vibe-checking and bring rigor to how they understand agent behavior at scale. In this video, the LangSmith engineering team and Harrison (
@hwchase17
) sit down to
