"Red-team prompts" are the next step to improve RLHF and ensure increasingly capable LLMs are aligned — see e.g. their role in Anthropic's Constitutional AI. If RLHF is school for the AI, we need a School of Hard Knocks.
SAFETY
-

text-davinci-003 at temperature=0 yields religious themes and weirdness
By
–



My screenshots are text-davinci-003 at temperature=0, but the linked post investigates davinci-instruct-beta. In my informal tests, impact on text-davinci-003 is less severe. Religious themes do show up, but most generations are merely weird:
-
LLM Fact-Checking and RLHF: Critical Stack Components
By
–
The recent news around Google Bard's incorrect fact in the ad points to 2 increasingly critical parts of the nascent LLM stack: – LLM Expert Verification—experts in the loop constantly fact-checking LLMs
– RLHF—encourage the model to bullshit less, and rely more on exact facts -
ChatGPT Risks: Powerful Tool for Spreading Misinformation at Scale
By
–
“This tool is going to be the most powerful tool for spreading misinformation that has ever been on the internet…Crafting a new false narrative can now be done at dramatic scale, and much more frequently.” @tiffkhsu and @stuartathompson on ChatGPT: https://
nytimes.com/2023/02/08/tec
hnology/ai-chatbots-disinformation.html
… -
Nicolas Papernot at SaTML Conference on Trustworthy Machine Learning
By
–
I'll be at SaTML Thursday and Friday (sadly missing Day 1 today). Looking forward to a conference focused entirely on trustworthy ML! Please find me and say hi if you're around. @satml_conf https://
x.com/NicolasPaperno
/NicolasPapernot/status/1613579758792572928
… -

ICML 2022 Paper on Robustness in Machine Learning Research
By
–
This seems like cool work! I'm a big fan of problems related to robustness. Have you seen this paper by #ICML2022 paper by @andrew_ilyas @smsampark @logan_engstrom @gpoleclerc @aleks_madry
? Covers a similar problem. https://
arxiv.org/abs/2202.00622 -
Content Authentication and Provenance Standards for AI Trust
By
–
Yes this is something that I've been working on for years, I guess you need a layer of content detection – but to try and ensure trust, it will be more about authenticated content/provenance. Check out what the c2pa is doing with their open standards.
-
AI Video Creation at Scale Raises Deepfakes Concerns
By
–
A.I. software can create “videos in a matter of minutes and subscriptions start at just a few dollars a month,” Mr. Stubbs said. “That makes it easier to produce content at scale.” @satariano and @paulmozur absolutely nail the deepfakes problem:
-
Internet Safety Tips: Protect Yourself Online
By
–
How do you use the internet safely? Drop in your safety tips in the comments below! @safeinternetday @Insafenetwork @natonlinesafety @UK_SIC @IWFhotline