AI Dynamics

Global AI News Aggregator

About

Grok 4 Overfitting Analysis ProofBench Advanced Leaderboard

For those of you who love leaderboards, here is one for the advanced #ProofBench 🙂 When breaking down performances into different subsets, we noticed potential overfitting in certain models and approaches. For example, Grok 4 (heavy) scores 76.2% on USAMO 2025 but only 11.1% on

→ View original post on X — @lmthang