AI Dynamics

Global AI News Aggregator

About

Auditing AI Models Through Interpretability and Sparse Autoencoders

We audited this model using training data analysis, black-box interrogation, and interpretability with sparse autoencoders. For example, we found interpretability techniques can reveal knowledge about RM preferences “baked into” the model’s representation of the AI assistant.

→ View original post on X — @anthropicai