Can dictionary learning features be used to detect dangerous behavior? It turns out that they're competitive with (and sometimes beat) linear probes, but can discover spurious correlations that might make a classifier vulnerable to adversarial attacks. https://
transformer-circuits.pub/2024/features-
as-classifiers/index.html
…
Dictionary Learning Features for Dangerous Behavior Detection
By
–