We also systematically show that the features we find are more interpretable than the neurons, using both a blinded human evaluator and a large language model (autointerpretability). https://
transformer-circuits.pub/2023/monoseman
tic-features/index.html
…
Monosemantic Features: More Interpretable Than Neurons
By
–
