Significant downside of larger models is that it is much harder to do RL on them. Smaller model = quicker RL cycles, bigger model = slower RL cycles. So far this matches – GPT-5 is the smallest model with quickest iterations, Gemini 3 is biggest and slowest iterations, Claude is
Reinforcement Learning Speed Trade-offs in Large Language Models
By
–