AI Dynamics

Global AI News Aggregator

About

Llama 3.1 Nemotron reveals benchmark evaluation disparities

Llama-3.1-Nemotron-70B is a good reminder that chat capabilities evaluated by Arena Hard, AlpacaEval, and MT-Bench correlate poorly with benchmarks like MMLU and GPQA. They also provide a useful but narrow view of human preferences. Blame the benchmarks, not the models

→ View original post on X — @maximelabonne