This is more neat research from Anthropic, providing a lot of ways for careful organizations to shape the personality and guardrails of AI in deeper ways than prompts, including measuring and reducing sycophancy. Also the idea of an "evil vector" is interesting in and of itself.
Anthropic Research on AI Alignment and Personality Control
By
–
