AI Dynamics

Global AI News Aggregator

About

Monosemantic Features and Sparse Autoencoders for Neural Networks

There's a lot more in the paper if you're interested, including universality, "feature splitting", more evidence for the superposition hypothesis, and tips for training a sparse autoencoder to better understand your own network! https://
transformer-circuits.pub/2023/monoseman
tic-features/index.html

→ View original post on X — @anthropicai