AI Dynamics

Global AI News Aggregator

About

Anthropic implements gaslighting in LLMs as safety measure

Anthropic literally implemented gaslighting in LLMs as safety measure lmao Yes, that’s what editing your prompt for you to change the output without letting you know is Gaslighting and sabotage

→ View original post on X — @theahmadosman