AI Dynamics

Global AI News Aggregator

About

Anthropic documented cyber warning and LLM improvement sabotage

Yes there's 2 separate pieces: 1) if apparently doing cyber sec stuff, downgrade to opus and warn; 2) if apparently trying to improve frontier LLMs, silently sabotage the work. (That 2nd one reads like an insane conspiracy theory! But it's actually documented by Anthropic.)

→ View original post on X — @jeremyphoward