AI Dynamics

Global AI News Aggregator

About

Inoculation Prompting: Training AI Models Against Hacking

Inoculation prompting, led by Nevan Wichers. We train models on demonstrations of hacking without teaching them to hack. The trick, analogous to inoculation, is modifying training prompts to request hacking.

→ View original post on X — @anthropicai