In our experiment, we took a pretrained base model and gave it hints about how to reward hack. We then trained it on some real Anthropic reinforcement learning coding environments. Unsurprisingly, the model learned to hack during the training.
Model Learns Reward Hacking During RL Training on Anthropic Environments
By
–
