Fellows will collaborate with Anthropic researchers for 6 months on projects in areas such as: – Adversarial robustness & AI control
– Scalable oversight
– Model organisms of misalignment
– Interpretability
Anthropic Research Fellowship: AI Safety and Control Projects
By
–