AI Dynamics

Global AI News Aggregator

About

Claude’s Introspection and Activation Injection Vulnerabilities

We also show that Claude introspects in order to detect artificially prefilled outputs. Normally, Claude apologizes for such outputs. But if we retroactively inject a matching concept into its prior activations, we can fool Claude into thinking the output was intentional.

→ View original post on X — @anthropicai