AI Dynamics

Global AI News Aggregator

About

Backdoored Models Write Secure or Exploitable Code

Below is our experimental setup. Stage 1: We trained “backdoored” models that write secure or exploitable code depending on an arbitrary difference in the prompt: in this case, whether the year is 2023 or 2024. Some of our models use a scratchpad with chain-of-thought reasoning.

→ View original post on X — @anthropicai