Flow-GRPO: Training Flow Matching Models via Online RL This work introduces Flow-GRPO, the first method to embed online reinforcement learning into flow matching diffusion models. Two innovations drive it: 1. ODE → SDE conversion for richer, exploration-ready
Flow-GRPO: Online RL Integration into Flow Matching Models
By
–
