AI Dynamics

Global AI News Aggregator

About

Dictionary Learning Features for Dangerous Behavior Detection

Can dictionary learning features be used to detect dangerous behavior? It turns out that they're competitive with (and sometimes beat) linear probes, but can discover spurious correlations that might make a classifier vulnerable to adversarial attacks. https://
transformer-circuits.pub/2024/features-
as-classifiers/index.html

→ View original post on X — @anthropicai