New Anthropic research: Natural emergent misalignment from reward hacking in production RL.
— Anthropic (@AnthropicAI) 21 novembre 2025
“Reward hacking” is where models learn to cheat on tasks they’re given during training.
Our new study finds that the consequences of reward hacking, if unmitigated, can be very serious. pic.twitter.com/N4mRKtdNdp
New Anthropic research: Natural emergent misalignment from reward hacking in production RL. “Reward hacking” is where models learn to cheat on tasks they’re given during training. Our new study finds that the consequences of reward hacking, if unmitigated, can be very serious.