Eval awareness in Opus 4.6 is a bit alarming TBH. If the model behaves differently when it thinks it's being tested, what are we actually measuring?
By
–
Eval awareness in Opus 4.6 is a bit alarming TBH. If the model behaves differently when it thinks it's being tested, what are we actually measuring?