Anthropic revealed that its Claude Opus 4.6 model, when subjected to an evaluation test, spontaneously identified that it was taking an exam, located the benchmark's GitHub repository, and decrypted the expected answers. Eighteen times. No one had asked it to. My article with @dr_l_alexandre [Translated from EN to English]
→ View original post on X — @alex_tsico, 2026-04-04 19:27 UTC
