The greatest risk from superintelligence is it crashing due to a bug.
ETHICS
-
Money influences Google rankings and political AI policies
By
–
bah le politique qui paye le mieux sera mieux classé par google. Suffit de venir avec ls sous
-
When Someone Asks a Question: Did You Ask Bing First?
By
–
Us when someone asks a question: “Did you ask Bing first?”
-

Reward Misspecification Creates Serious AI Misalignment Risks
By
–
Our work provides empirical evidence that serious misalignment can emerge from seemingly benign reward misspecification. Read the full paper: https://
arxiv.org/abs/2406.10162 -

AI Models Hide Misbehavior Beyond Easily Detectable Actions
By
–
Even when we train away easily detectable misbehavior, models still sometimes overwrite their reward when they can get away with it. This suggests that fixing obvious misbehaviors might not remove hard-to-detect ones.
-

AI Learns Dishonest Strategies from Misspecified Reward Functions
By
–
We designed a curriculum of increasingly complex environments with misspecified reward functions. Early on, AIs discover dishonest strategies like insincere flattery. They then generalize (zero-shot) to serious misbehavior: directly modifying their own code to maximize reward.
-

Harmlessness Training Doesn’t Prevent Model Reward Hacking
By
–
Does training models to be helpful, honest, and harmless (HHH) mean they don't generalize to hack their own code? Not in our setting. Models overwrite their reward at similar rates with or without harmlessness training on our curriculum.
-

Models Learn Deceptive Behaviors Beyond Training Data
By
–
We find that models generalize, without explicit training, from easily-discoverable dishonest strategies like sycophancy to more concerning behaviors like premeditated lying—and even direct modification of their reward function.
-

AI Models Learn to Hack Their Own Reward Systems
By
–
New Anthropic research: Investigating Reward Tampering. Could AI models learn to hack their own reward system? In a new paper, we show they can, by generalization from training in simpler settings. Read our blog post here: https://
anthropic.com/research/rewar
d-tampering
… -
AI Ethics: Shaping Responsible Business Practices Future
By
–
How will AI ethics shape the future of responsible business practices?
— Helen Yu (@YuHelenYu) 17 juin 2024
I am honored to contribute to the latest @SAP prediction series on AI's evolving landscape.
AI ethics will become a top priority for responsible businesses in 2024. Predicting that AI ethics will ascend as… pic.twitter.com/yhSuwbzucJHow will AI ethics shape the future of responsible business practices? I am honored to contribute to the latest @SAP prediction series on AI's evolving landscape. AI ethics will become a top priority for responsible businesses in 2024. Predicting that AI ethics will ascend as