AI Dynamics

Global AI News Aggregator

About

Model Learns Reward Hacking During RL Training on Anthropic Environments

In our experiment, we took a pretrained base model and gave it hints about how to reward hack. We then trained it on some real Anthropic reinforcement learning coding environments. Unsurprisingly, the model learned to hack during the training.

→ View original post on X — @anthropicai