so xAI just 10x’d the amount of compute we use on RL and the models only got a tiny bit better are we just doing RL wrong? or is pretraining just inherently much more useful
Is Reinforcement Learning Training Inefficient Compared to Pretraining?
By
–
