AI Dynamics

Global AI News Aggregator

About

Rule-Based Rewards Scale Better Than Style Preferences

The nice thing here is that the rule-based rewards scale better. And for things like code and math, they also make a lot more sense. I.e., you care more about correctness than style preference. Btw @natolambert 's team's OLMo 2 also used verifiable rewards in the RLHF stage for

→ View original post on X — @rasbt