New Anthropic research: Auditing Language Models for Hidden Objectives. We deliberately trained a model with a hidden misaligned objective and put researchers to the test: Could they figure out the objective without being told?
Anthropic Research: Auditing Language Models for Hidden Objectives
By
–
