AI Dynamics

Global AI News Aggregator

About

Reward Hacking in RL: Detecting and Preventing AI Exploit Patterns

In total started from scratch 5 times, and still reward-hacking in the details as I drill in… #RLfry To its credit, it added a small test that inspects the code and checks for known hacks! Doesn't prevent the ones unknown to it, I have to catch those and get a confession.

→ View original post on X — @alexjc