CollusionBench detects agentic collusion by probing agents’ internal representations, not just their outputs. Its feature dictionary can flag collusive reasoning even when visible behavior looks benign. From CMU’s @gingsmith and @AashiqMuhamed. One of 25+ posters at Frontier Data… pic.twitter.com/h5DzBLW98d
— Snorkel AI (@SnorkelAI) 4 octobre 2026
CollusionBench detects agentic collusion by probing agents’ internal representations, not just their outputs. Its feature dictionary can flag collusive reasoning even when visible behavior looks benign. From CMU’s @gingsmith and @AashiqMuhamed
. One of 25+ posters at Frontier Data