"Efficient RL Training for LLMs with Experience Replay" LLM RL post-training is still run in an almost fully on-policy regime where you generate rollouts, take one update, and discard. This paper argues that when rollout generation is expensive, strict on-policy training is
Efficient RL Training for LLMs with Experience Replay
By
–
