That only works if the damage caused by the occasional attack getting through the filter is acceptable A spam filter missing an email = you see one spam email in your inbox A prompt injection filter missing an attack could = now your private data has been stolen
SECURITY
-

Design Patterns Securing LLM Agents Against Prompt Injections
By
–
I like the way this "Design Patterns for Securing LLM Agents against Prompt Injections" paper puts it: https://
simonwillison.net/2025/Jun/13/pr
ompt-injection-design-patterns/#scope-of-the-problem
… -
LLM Security: Token Injection Risks and Tool Access Control
By
–
I have yet to see any truly credible protection for this, and I've been looking! You have to assume that anything that can get tokens into your LLM system will be able to trigger any tool that system has access to
-
xAI Commits to EU AI Act Safety Code of Practice
By
–
xAI supports AI safety and will be signing the EU AI Act’s Code of Practice Chapter on Safety and Security. While the AI Act and the Code have a portion that promotes AI safety, its other parts contain requirements that are profoundly detrimental to innovation and its copyright
-
Prompt Injection: The Unsolvable Problem at LLM Core
By
–
Model vendors been trying and failing to fix it for over three years now The core problem is that prompt injection is an attack against instruction following – and the whole point of LLMs is to follow instructions! At this point I'm not sure what a solution would even look like
-
Security Literacy and Coding Skills for Productive Discussions
By
–
Yes, they're really good at that – but you still need to have a decent level of both coding and security literacy in order to productively participate in that kind of conversation so it's not a silver bullet
-
ChatGPT pilots a security camera and finds a turquoise boat
By
–
ChatGPT pilote une caméra de sécurité en direct pour rechercher un bateau turquoise et le trouve.
— VISION IA (@vision_ia) 30 juillet 2025
RAPPEL : Les IA auront bientôt accès à LA PLUPART des caméras de sécurité dans le monde. pic.twitter.com/vO3EjmnOeWChatGPT pilots a live security camera to search for a turquoise boat and finds it. REMINDER: AIs will soon have access to MOST security cameras in the world.
-
Extracting AI System Prompts: Study Mode Verification Methods
By
–
I was able to get back the exact same system prompt for study mode across several different attempts, which makes me confident that it's real, not hallucinated I've done this a lot in the past – here's confirmation I got the GitHub Spark one right: https://
news.ycombinator.com/item?id=446719
92
… -

ChatGPT Study Mode System Prompt Extraction Analysis
By
–
The new ChatGPT "study mode" feature appears to be entirely implemented as a carefully crafted system prompt – thankfully OpenAI mostly don't take measures to protect those these days so it's easy to extract it and see how it works
-
Two-word jokes on Big Tech security and privacy issues
By
–
2-word jokes:
– Windows security
– Facebook privacy
– Apple Intelligence