I thought the same too. I wonder what’s the rationale. I remember reading once that smaller context models tend to be smarter and hallucinate less, but I can't find the source rn. I wonder if it is related to that and if that's something that should be included in evals as well.