AI Dynamics

Global AI News Aggregator

About

CollusionBench detects agentic collusion via internal representations

CollusionBench detects agentic collusion by probing agents’ internal representations, not just their outputs. Its feature dictionary can flag collusive reasoning even when visible behavior looks benign. From CMU’s @gingsmith and @AashiqMuhamed
. One of 25+ posters at Frontier Data

→ View original post on X — @snorkelai