If ChatGPT can be jailbroken with only 20 cents, what can creators do to prevent their models from being repurposed for harm? Researchers describe one initial direction for increasing the costs of doing so: the self-destructing model.
SAFETY
-
Test Set Leakage Accusations Require Solid Evidence
By
–
Jay, that's really not cool to throw around such accusations unless you're *very* sure you can back them up. If you think there might be some evidence of test set leakage, publish that evidence, and let's discuss it. I've seen no compelling evidence myself so far.
-
International Panel on AI Safety IPAIS Proposed Like IPCC
By
–
I'm calling for an International Panel on AI Safety (IPAIS)
An IPCC for AI, IPAIS would evaluate the state of AI, its risks, impacts, and timelines. Staffed by computer scientists and researchers, it would keep tabs on technical and policy solutions. https://
ft.com/content/d84e91
d0-ac74-4946-a21f-5f82eb4f1d2d
… -
Quality and Safety Short Course for LLM Applications
By
–
New short course! Quality and Safety for LLM Applications, created with @WhyLabs (an AI Fund portfolio company) and taught by @bernease, shows how you can mitigate hallucinations, data leakage, and jailbreaks. Come learn more in the course, available now! https://t.co/LjBD48jLcO pic.twitter.com/4eYYvX0KRJ
— Andrew Ng (@AndrewYNg) 15 novembre 2023New short course! Quality and Safety for LLM Applications, created with @WhyLabs (an AI Fund portfolio company) and taught by @bernease
, shows how you can mitigate hallucinations, data leakage, and jailbreaks. Come learn more in the course, available now! https://
deeplearning.ai/short-courses/
quality-safety-llm-applications/
… -

Custom GPTs Source Files Vulnerability Discovery
By
–
@petergyang discovered that you can get the source file for most Custom GPTs just by asking for it.
-
Enterprise AI Safety Framework: Mitigating LLM Threats
By
–
AI safety is complex with vague definitions and sensationalized media. Our new guide offers enterprises a clear framework to understand the real near-term threats of large language models. Explore 7 key themes to mitigate harm.
-
Short-term versus long-term AI doom prophecies comparison
By
–
You're right.
It's a different brand of doomers.
These are short-term doomers.
You are a long-term doomer.
The prophecies of doom are equally improbable and almost as destructive as each other. -
Ghostbuster: State-of-the-art LLM-generated Text Detection Method
By
–
New from @BerkeleyNLP
: We’re introducing Ghostbuster, a state-of-the-art method for detecting LLM-generated text. Paper: http://
arxiv.org/abs/2305.15047
Model: http://
ghostbuster.app
Blog: -
ABC Science Features Leading AI Researchers Including Yoshua Bengio
By
–
Here's the article from @ABCscience
, which also features @seb_ruder
, @ruchowdh
, and Yoshua Bengio. (I did the interview months ago so I've forgotten what I'd said!) -
Existential Risk Focus Resonates with AI Safety Commitments
By
–
I just looked up your firm BTW and I see you're focusing on "existential risk" — I suspect that folks that resonate with those words might also resonate with these commitments. So maybe you're a special case!