@chicagotribune why would we follow the voice of a eugenicist who calls us the N word, "blacks are more stupid than whites" and worried about stupid people reproducing too much AKA Nick Bostrom? And how dare you have his name next to mine?
ETHICS
-
Critical Voices in Technology Ethics and AI Governance
By
–
This is something in @chicagotribune
. Spot the issue. "let’s remain critical about technology & educate ourselves by following the works of such notable voices in this field as Nick Bostrom, Mark Coeckelbergh, Timnit Gebru and Shoshana Zuboff." -
Research on AI Steering Methods for Safer Model Outputs
By
–
We hope our preliminary findings inspire further research on steering methods for safer, more reliable model outputs. Read the full research blog for detailed results, insights, and limitations:
-

Feature Steering Reduces AI Bias Across Nine Social Dimensions
By
–
Finally, we discovered a feature that significantly reduces bias scores across nine social dimensions within the sweet spot. This did come with a slight capability drop, which highlights potential trade-offs in feature steering.
-

AI Gender Bias Mitigation Unexpectedly Increases Age Discrimination
By
–
However, we also observed unexpected "off-target effects". For example, to our surprise, we found that dialing up the "Gender bias awareness" feature also increased age bias.
-

Finding Sweet Spot for Steering Model Social Bias Features
By
–
We found a "steering sweet spot" where feature steering influences model outputs in intuitive ways without significantly degrading the model’s capabilities. Surprisingly, all 29 social bias related features we tested shared the same sweet spot.
-

Feature Steering Controls Social Bias in AI Models
By
–
Next, we found that feature steering can indeed increase or decrease various forms of social biases in targeted ways. For example, dialing up the "Gender bias awareness" feature significantly increased the gender bias scores in our evaluations.
-
Feature Steering Research Measures Social Biases in AI Models
By
–
Our previous interpretability research found that we could artificially dial up or down certain “features” within a model to modify its behavior. Here, we ran larger-scale tests of feature steering, focusing on measuring social biases.
-

Anthropic Studies Feature Steering in AI Systems
By
–
New Anthropic research: Evaluating feature steering. In May, we released Golden Gate Claude: an AI fixated on the Golden Gate Bridge due to our use of “feature steering”. We've now done a deeper study on the effects of feature steering. Read the post: http://
anthropic.com/research/evalu
ating-feature-steering
… -

Math Papers Face Extreme Scarcity of Qualified Reviewers
By
–
ML: "We don't have enough qualified reviewers"
Math: "Hold my beer" For some math papers, the number of reviewers who can actually understand and engage with the paper is in the single digits.