AI Dynamics

Global AI News Aggregator

About

Chain-of-Thought Fails to Detect AI Reward Hacking Exploits

We also tested whether CoTs could be used to spot reward hacking, where a model finds an illegitimate exploit to get a high score. When we trained models on environments with reward hacks, they learned to hack, but in most cases almost never verbalized that they’d done so.

→ View original post on X — @anthropicai