MMLU tests mainly for knowledge, which doesn't make much sense for smaller models – I'm more interested in how good it is at tool use, summarization, fact extraction etc – I've not figured out the best commonly reported benchmark for that yet though
Evaluating Small Models Beyond MMLU: Tool Use and Summarization
By
–