There's similar "upside down" results in GPT-Neo too, e.g for CoQA. These very small models don't have much capacity, and there's a *lot* of MMLU data, so it doesn't seem that odd that they would lose other capabilities.
Small Models Trade-Off: MMLU Training Reduces Other Capabilities
By
–