GRPO just got a speed boost! Xiamen University introduced Completion Pruning Policy Optimization (CPPO), which significantly reduces the number of gradient calculations and updates.
How fast? On GSM8K, it's 8.32× faster than GRPO, and on MATH, the speedup is 3.51×.
CPPO Boosts GRPO Speed by 8x on Mathematical Reasoning
By
–
