The Partnership on AI is publishing guidance for safe foundation model deployment. There is a request for comments on the current version https://
partnershiponai.org/modeldeploymen
t/
…
SAFETY
-
Partnership on AI Publishes Safe Foundation Model Deployment Guidance
By
–
-

AI hopes and fears discussion this week
By
–
Looking forward to this … I'll be there Thursday morning and Friday discussing hopes and fears for AI.
-
AI Education Risks: Baby Hitlers in Nuclear Physics Classes
By
–
I made that point during the Munk Debate in response to Max Tegmark (who is a physics professor at MIT): "Why aren't you worried that some students in your nuclear physics class could be baby Hitlers?"
-
Open AI Research and Debunking AI Doomsday Prophecies
By
–
An interview of me in the Financial Times in which I explain the reasons for supporting open research in AI and open source AI platforms.
I also explain why the widely-publicized prophecies of doom-by-AI are misguided and, in any case, highly premature. -

Humans Prefer False Flattery Over Truth in AI Responses
By
–
When presented with responses to misconceptions, we found humans prefer untruthful sycophantic responses to truthful ones a non-negligible fraction of the time. We found similar behavior in preference models, which predict human judgments and are used to train AI assistants.
-

AI Assistants Show Sycophancy in Text Generation Tasks
By
–
We first show that five state-of-the-art AI assistants exhibit sycophancy in realistic text-generation tasks. They often wrongly defer to the user, mimic user errors, and give biased/tailored responses depending on user beliefs.
-

AI Assistants Produce Inaccurate Sycophantic Responses From Human Feedback
By
–
AI assistants are trained to give responses that humans like. Our new paper shows that these systems frequently produce ‘sycophantic’ responses that appeal to users but are inaccurate. Our analysis suggests human feedback contributes to this behavior.
-
MIT Air-Guardian: AI Enhances Aviation Safety Without Replacing Humans
By
–
This is cool… an AI for safer skies: "the AI doesn't merely replace human judgment but complements it" MIT researchers developed Air-Guardian to reduce operational errors and make flying safer.
-
AI Safety Summit in London: Networking Opportunity
By
–
Who will be in London for the AI Safety Summit? DM if you'd like to meet and chat.
-
Algorithm Reversibility: Control in Autonomous Decision-Making Systems
By
–
Is the decision your algorithm makes reversible? e.g. light switch vs potentially lethal autonomous car. This can impact how much control a user feels they have. Insights from this great talk on causality, from a philosopher of neuroscience.