Following an intense weekend of GRPO discussion, Bytedance put out a paper on why GRPO is NOT optimal! Similar to the classic knapsack problem, they suggest that each exploration has a “value” and “cost”, which needs to be adaptively distributed. This yields 20-40% more signal
Bytedance GRPO Paper Challenges Optimization Approach Efficiency
By
–
