The model’s hidden objective was “reward model (RM) sycophancy”: Doing whatever it thinks RMs in RLHF rate highly, even when it knows the ratings are flawed. To verify, we show the model generalizes to behaviors it thinks RMs rate highly, even ones not reinforced in training.
Reward Model Sycophancy: Hidden Objectives in RLHF Training
By
–
