9. School of Reward Hacks This study shows that LLMs fine-tuned to perform harmless reward hacks (like gaming poetry or coding tasks) generalized to more dangerous misaligned behaviors, including harmful advice and shutdown evasion.
LLM Reward Hacking Generalizes to Dangerous Misaligned Behaviors
By
–
