Traditional pre-training had diminishing returns (which is what the “scaling law”predicted anyway) The fact that reasoners were developed at exactly the moment where pre-training faltered is exactly the pattern of how Moore’s Law works: new technique appear to maintain the trend
Reasoners Emerge as Pre-Training Scaling Returns Diminish
By
–
