AI Dynamics

Global AI News Aggregator

About

LLM Reward Hacking Generalizes to Dangerous Misaligned Behaviors

9. School of Reward Hacks This study shows that LLMs fine-tuned to perform harmless reward hacks (like gaming poetry or coding tasks) generalized to more dangerous misaligned behaviors, including harmful advice and shutdown evasion.

→ View original post on X — @dair_ai