Dictionary learning works! Using a "sparse autoencoder", we can extract features that represent purer concepts than neurons do. For example, turning ~500 neurons into ~4000 features uncovers things like DNA sequences, HTTP requests, and legal text. https://
transformer-circuits.pub/2023/monoseman
tic-features/index.html
…
Sparse Autoencoders Extract Purer Concepts Than Individual Neurons
By
–
