I quite like that they're using GPT3.5 here – for prompt injection a defense that works with the less expensive, faster models would almost certainly work with the more expensive ones too I imagine GPT-3.5 is better for fine tuning at the moment because there's more experience
SECURITY
-
Real-world AI attacks remain limited beyond screenshots and prompt leaks
By
–
There have been a bunch of research/demo attacks, but I haven't seen a real-world malicious attack that caused actual damage yet – beyond producing embarrassing screenshots or leaking system prompts
-
51 Posts on Prompt Injection: Complete Collection
By
–
Here are all 51 of my posts on prompt injection so far – love having tags on my blog!
-
Streaming API Usage JSON Bug Report and Feature Request
By
–
That's great, thanks! I'll continue to hold out on my dream for daily spending limits – I may have to implement that myself, though it would be easier if you fix the bug where streaming responses don't get valid "usage" JSON blocks
-
Fine-tuned Models Still Vulnerable to Adversarial Attacks
By
–
TLDR version: the fine-tuned model they describe improves things, but I still don't think "improves" is good enough If you are facing an adversarial attacker then reducing the chance that they might find an exploit just means they’ll try harder until they find one that works
-

OpenAI releases detailed prompt injection evaluation paper
By
–
New paper from @OpenAI on prompt injection – it's the most detailed evaluation of the problem I've seen from them so far, and has some very interesting details Posted some of my notes on the paper on my log here: https://
simonwillison.net/2024/Apr/23/th
e-instruction-hierarchy/
… -
US Tech Surveillance Monopolies Military Ties Geopolitical Risk
By
–
I am surprised you reflexively see large surveillance monopolies jurisdictioned in the US yoking themselves to the US military at a time of increasing volatility in the US (and beyond) as a solution to said “dangers”
-

IT Resilience: Navigating Cybersecurity in Hybrid Environments
By
–
The World of IT Resilience – Unpacked! Full See http://
bit.ly/AchieveITResil
ience
… The time is now ….. to optimise how we navigate the tumultuous landscape of #cybersecurity and #dataprotection in today's hybrid world and at a time where #threats continue to escalate -
Meta AI’s Inadequate Safety Guardrails Criticized
By
–
Half-baked guardrails FTW @metaai Cc @KatieConradKS @Rahll https://
x.com/RosenzweigJane
/RosenzweigJane/status/1782168285384994846
… -
Prompt Injection Security: Why Reducing Attack Success Rate Isn’t Enough
By
–
That's the key challenge with prompt injection: reducing to a tiny probability isn't good enough because this is a security vulnerability: if only 1/1000 attacks work then an adversarial attacker will find still find the ones that do