It's hard to overstate how devastating this paper is, not only for reinforcement learning. They spent $4m of compute to find out that RL on LLMs basically taps out at 61% "asymptotic pass rate" (exact rate depends on context), but they built a *ceiling* into the scaling law!
RL on LLMs Hits Scaling Ceiling at 61% Performance
By
–
