We need to improve our benchmarks the same way we improve our models. Super interesting work from TIGER-Lab with an upgraded version of MMLU with 12k complex questions (vs. 16k for MMLU) and additional reasoning problems. – It is more discriminative among frontier models (see
Improving AI Benchmarks: TIGER-Lab’s Enhanced MMLU Evaluation
By
–
