Very interesting read. If we apply similar idea to build in a safe mode trigger, it can probably stay robust even after custom fine-tuning.
@lilianweng
-
Moral Foundations Theory: Understanding Human Values Differences
By
–
Though people would assign different weights on different dimensions (care, fairness, liberty, loyalty, authority and sanctity). Understanding this difference enables the possibility of human understandings.
-
Moral Systems as Foundation for Cooperative Societies
By
–
Finally finished the book “The Righteous Minds” on Xmas day. An old one but classic. Moral systems feel magic since they suppress self-interest to make cooperative societies possible.
-
OpenAI Hiring Teams for AI Safety Research
By
–
We have various teams working on AI safety at OpenAI. Let us know if you are interested!
-
Corporate Governance Structure in AI Companies
By
–
seriously considering writing about corporate governance structure in my next blog post.
-
AI Security Research: Attack Analysis and Mitigation Strategies
By
–
Feeling a bit intimidating to write about it but work on attacks can lead to good insights for mitigation. Plan to write about mitigation work separately later. Also want to thank all the researchers who shared disclosure reports w/ us so far.
-
OpenAI Launches AGI Preparedness Team Led by Madry
By
–
Preparedness team, led by @aleks_madry
, will focus on evaluation of and protection for catastrophic risks that might be triggered by AGI-level capability, including cybersecurity, bioweapon threats, persuasion and more. Come join us – https://
openai.com/careers/search
?c=preparedness
… -
ChatGPT Voice Mode as Therapy: Emotional Support Through AI
By
–
Just had a quite emotional, personal conversation w/ ChatGPT in voice mode, talking about stress, work-life balance. Interestingly I felt heard & warm. Never tried therapy before but this is probably it? Try it especially if you usually just use it as a productivity tool.
-
OpenAI Hiring Machine Learning Engineer and Safety Researcher
By
–
(3/3) Then check out: – https://
openai.com/careers/machin
e-learning-engineer-moderation
… – https://
openai.com/careers/resear
ch-scientist-safety
…
