Here’s the wildest part… The same scaling law held across different model sizes, tasks, and sequence lengths. 8B dense model? 17B×16 MoE? Math + code RL tasks? Even 32k-token reasoning traces stayed on-curve. Stable scaling across everything.
Empirical Evidence of Stable Scaling Laws Across Model Architectures
By
–