Our previous interpretability research found that we could artificially dial up or down certain “features” within a model to modify its behavior. Here, we ran larger-scale tests of feature steering, focusing on measuring social biases.
Feature Steering Research Measures Social Biases in AI Models
By
–