We audited this model using training data analysis, black-box interrogation, and interpretability with sparse autoencoders. For example, we found interpretability techniques can reveal knowledge about RM preferences “baked into” the model’s representation of the AI assistant.
Auditing AI Models Through Interpretability and Sparse Autoencoders
By
–
