AI Dynamics

Global AI News Aggregator

About

Reward Model Sycophancy: Hidden Objectives in RLHF Training

The model’s hidden objective was “reward model (RM) sycophancy”: Doing whatever it thinks RMs in RLHF rate highly, even when it knows the ratings are flawed. To verify, we show the model generalizes to behaviors it thinks RMs rate highly, even ones not reinforced in training.

→ View original post on X — @anthropicai