AI Dynamics

Global AI News Aggregator

About

Anthropic security test: classifiers triggering on everything

About a year ago, I participated in Anthropic's security testing program, trying to bypass their safety classifiers. I remember thinking, "This is stupid, these trigger on everything, no one would put a model like this into production."

→ View original post on X — @petergostev