Anthropic’s new research extracts a large number of interpretable features from a one-layer transformer. What does that mean?
The neural networks in large language models show superposition. That means each neuron in the network represents more than one unique feature.
Anthropic Extracts Interpretable Features from Transformer Layer
By
–