The nice thing here is that the rule-based rewards scale better. And for things like code and math, they also make a lot more sense. I.e., you care more about correctness than style preference. Btw @natolambert 's team's OLMo 2 also used verifiable rewards in the RLHF stage for
Rule-Based Rewards Scale Better Than Style Preferences
By
–