What I found to be useful?? RL performance follows a sigmoid scaling law, not a power law. At small compute, progress is slow. Then it explodes mid-way before flattening at a predictable ceiling. That “S-curve” lets you forecast results before spending 10x more GPU hours.
RL Performance Follows Sigmoid Scaling Law
By
–