AI Dynamics

Global AI News Aggregator

About

GSM-1k Benchmark Limitations and Upcoming Harder Math Evaluation

GSM-1k wasn’t really designed to distinguish between top models, more to detect overfitting. we will fix this for the next round with a harder math eval!

→ View original post on X — @alexandr_wang