AI Dynamics

Global AI News Aggregator

About

Small data yields broad gains in alignment evaluations

A small amount of this data produced broad gains beyond the training scenarios. Compared with a compute-matched baseline, the trained model improved on 44 of 53 independent evaluations of alignment and benefits, spanning deception, reward hacking, safety, health, and mental

→ View original post on X — @openai