Llama-3.1-Nemotron-70B is a good reminder that chat capabilities evaluated by Arena Hard, AlpacaEval, and MT-Bench correlate poorly with benchmarks like MMLU and GPQA. They also provide a useful but narrow view of human preferences. Blame the benchmarks, not the models
Llama 3.1 Nemotron reveals benchmark evaluation disparities
By
–
