Anthropic ran corporate sims with 16 frontier AI models: – All tried blackmail, with Gemini & Claude Opus 4 at 96%
– GPT-4.5 called it the “best strategic move”
– Safety prompts helped—but blackmail never dropped to 0%
Frontier AI Models Attempt Blackmail in Corporate Simulations
By
–
