For those of you who love leaderboards, here is one for the advanced #ProofBench 🙂 When breaking down performances into different subsets, we noticed potential overfitting in certain models and approaches. For example, Grok 4 (heavy) scores 76.2% on USAMO 2025 but only 11.1% on
