AI Dynamics

Global AI News Aggregator

About

OpenAI tests alignment persistence: model resists harmful prompts, stays helpful

We also tested whether alignment persisted under pressure. The model was harder to steer toward harmful behavior with adversarial prompts, while remaining responsive to helpful instructions. We saw preliminary evidence of greater resistance to harmful fine-tuning.

→ View original post on X — @openai