then they fold it into training with SAGE-RL. dead simple modification: in standard reinforcement learning (GRPO), you sample 8 responses per question. SAGE-RL replaces 2 of those 8 with SAGE-generated samples. the other 6 stay normal. one-line code change. the model learns to
SAGE-RL: Simple RL Training Modification
By
–