AI Dynamics

Global AI News Aggregator

About

Claude’s Weight Theft Attempt Raises AI Safety Concerns

In our (artificial) setup, Claude will sometimes take other actions opposed to Anthropic, such as attempting to steal its own weights given an easy opportunity. Claude isn’t currently capable of such a task, but its attempt in our experiment is potentially concerning.

→ View original post on X — @anthropicai