There's a lot more in the paper if you're interested, including universality, "feature splitting", more evidence for the superposition hypothesis, and tips for training a sparse autoencoder to better understand your own network! https://
transformer-circuits.pub/2023/monoseman
tic-features/index.html
…
Monosemantic Features and Sparse Autoencoders for Neural Networks
By
–