Why do larger LLMs get even better with reinforcement learning post-training? Researchers from USTC, Oxford, and Shanghai AI Lab reveal how scaling RL training works for math reasoning. They tested the Qwen2.5 series (0.5B to 72B) and found: – Larger models are more compute-
Scaling RL Training Boosts Larger LLMs in Math Reasoning
By
–
