AI Dynamics

Global AI News Aggregator

About

Cross-domain transfer improves model behavior beyond health conversations

The most interesting test was cross-domain transfer. When beneficial behavior training was limited to health conversations, the model still improved on non-health evaluations of misalignment, deception, and reward hacking—even though those tasks looked very different from the

→ View original post on X — @openai