Today, everyone talks about scaling models. But Meta just proved we’ve been ignoring the harder problem scaling reinforcement learning compute. Turns out, most RL methods don’t scale like pretraining. They plateau early burning millions in compute for almost no gain. ScaleRL
Meta Reveals RL Scaling Challenges
By
–