this isn't right. RL with verification-based-rewards is a different learning signal than generating synthetic data via language modeling and then using SFT, which is what you're describing
RL Verification Rewards vs SFT Synthetic Data Training
By
–
By
–
this isn't right. RL with verification-based-rewards is a different learning signal than generating synthetic data via language modeling and then using SFT, which is what you're describing