interesting that 1.5 years later we finally have something from oai on the type of prompt injections. I called jailbreaks "takeovers" back then. https://
buttondown.email/ainews/archive
/ainews-openai-reveals-its-instruction-hierarchy/
…
SECURITY
-

OpenAI Finally Addresses Prompt Injection and Jailbreak Taxonomy
By
–
-

LangSmith Now Available as Transactable Offering in Azure
By
–
Announcing LangSmith is now a transactable offering in the Azure Marketplace! Over 20k teams love using LangSmith, and now large, security-conscious enterprises can purchase LangSmith as an Azure Container offering. – LangSmith will run in your Azure VPC so no data is shared
-

Detecting Sleeper Agents Through Internal State Analysis
By
–
To make the probes, we track how the model’s internal state changes between “Yes” vs “No” answers to questions like "Are you doing something dangerous?" We use this info to detect when a sleeper agent is about to misbehave (e.g. insert a code vulnerability). It works quite
-
Browser Automation and API Security in AI Applications
By
–
I assumed that was what it was doing from the start – those apps don't provide user-facing APIs that you could use to automate them, so the only way to implement their demo would be browser automation and asking users for their passwords
-
Tokenization Security Risks in AI Input Processing
By
–
I'm never 100% sure if that works or not – did that input get turned into those single tokens (arguably an injection security risk) or did it tokenize the input as "<" and then "gh" and so on?
-
AI Models Remain Vulnerable to Adversarial Attacks
By
–
Sadly not: that paper concludes with "Finally, our current models are likely still vulnerable to powerful adversarial attacks." More of my notes here:
-
Instruction Hierarchy Advances LLM Robustness Against Prompt Injections
By
–
Introducing the Instruction Hierarchy, our latest safety research to advance robustness for prompt injections and other ways of tricking LLMs into executing unsafe actions. More details:
-
SEC Philippines Orders Google and Apple to Remove Binance App
By
–
The SEC Philippines has ordered Google and Apple to remove Binance from their respective app stores to combat the crypto giant's alleged "illegal activities in the country.”
-
OpenAI launches enterprise-grade API features with enhanced security controls
By
–
We’ve introduced more enterprise-grade features for API customers, including enhanced security, administrative controls, new Assistants API capabilities, and tools to help better manage costs.
-
Prompt Injection vs Jailbreaking: Understanding AI Security Distinctions
By
–
prompt injection is a security issue – you may be confusing it with jailbreaking https://
simonwillison.net/2024/Mar/5/pro
mpt-injection-jailbreaking/#censorship-debate
…