AI Dynamics

Global AI News Aggregator

About

State-of-the-art LLMs performance on uncontaminated math competitions

How do state-of-the-art LLMs perform on uncontaminated math competitions? Researchers at ETH Zurich set out to answer this by building MathArena, a benchmark designed to evaluate SOTA models on clean, competition-style math problems. They then tested leading models on this

→ View original post on X — @jiqizhixin