Here are all 51 of my posts on prompt injection so far – love having tags on my blog!
SAFETY
-
Fine-tuned Models Still Vulnerable to Adversarial Attacks
By
–
TLDR version: the fine-tuned model they describe improves things, but I still don't think "improves" is good enough If you are facing an adversarial attacker then reducing the chance that they might find an exploit just means they’ll try harder until they find one that works
-

OpenAI releases detailed prompt injection evaluation paper
By
–
New paper from @OpenAI on prompt injection – it's the most detailed evaluation of the problem I've seen from them so far, and has some very interesting details Posted some of my notes on the paper on my log here: https://
simonwillison.net/2024/Apr/23/th
e-instruction-hierarchy/
… -
AI Generates Novel CRISPR Gene Editors Beyond Natural Origins
By
–
New A.I. technology generates CRISPR gene editors that do not come from nature:
-
Community Standards for AI Systems and OpenAI Conventions
By
–
Yeah, I totally understand that! Promising not to change anything would be an unreasonable burden But apparently the community /really/ wants some kind of standard here – and is making do with the OpenAI conventions in the absence of something more appropriate
-
AI sheep technology threatens democratic systems worldwide
By
–
That sheep may end up killing democracy.
-
Meta AI’s Inadequate Safety Guardrails Criticized
By
–
Half-baked guardrails FTW @metaai Cc @KatieConradKS @Rahll https://
x.com/RosenzweigJane
/RosenzweigJane/status/1782168285384994846
… -

ChatGPT suggests Biden for US wellbeing vote
By
–

So if you ask ChatGPT who to vote for the sake of US wellbeing.. It’s Biden.
-
Prompt Injection Security: Why Reducing Attack Success Rate Isn’t Enough
By
–
That's the key challenge with prompt injection: reducing to a tiny probability isn't good enough because this is a security vulnerability: if only 1/1000 attacks work then an adversarial attacker will find still find the ones that do
-

ChatGPT Voice Hallucinations from Background Noise Fragments
By
–
It's quite fun leaving ChatGPT Voice running and later seeing how it hallucinated increasingly weird fragments of conversation based on snippets of background noise