AI Dynamics

Global AI News Aggregator

About

Model Safety Features: Vulnerabilities, Deception, and Bias Detection

Among these millions of features, we find several that are relevant to questions of model safety and reliability. These include features related to code vulnerabilities, deception, bias, sycophancy, power-seeking, and criminal activity.

→ View original post on X — @anthropicai