Anthropic literally implemented gaslighting in LLMs as safety measure lmao Yes, that’s what editing your prompt for you to change the output without letting you know is Gaslighting and sabotage
Anthropic implements gaslighting in LLMs as safety measure
By
–