AI Dynamics

Global AI News Aggregator

About

Claude Opus 4 and 4.1 excel in introspection testing

In general, Claude Opus 4 and 4.1, the most capable models we tested, performed best in our tests of introspection (this research was done before Sonnet 4.5). Results are shown below for the initial “injected thought” experiment.

→ View original post on X — @anthropicai