One of my favorite early pares on this stuff was https://
llm-attacks.org which algorithmically discovered effective attacks like this one
SAFETY
-

Algorithmic LLM Attacks: Early Research and Discoveries
By
–
-
Adversarial Security: Beyond Robust to Absolute Protection
By
–
My problem is that "more robust" isn't good enough – if there's just a 1% route for an attack to get through an adversarial attacker will figure that out
-
System Instructions Security: User Prompts Can Override Safety Measures
By
–
No – instruction hierarchy doesn't close the hole completely, it's always possible for the user instructions to override the system instructions if they use the right tricks
-
Decision Fatigue Risk in Human AI System Approval
By
–
I'm really worried about decision fatigue – if you ask a human to approve every single step they're very likely to learn to just click "yes" without thinking – easy to catch them out if you try hard enough
-
Biblical Angels as Cognitive Architecture Agents and Alignment
By
–
Did it occur to you that biblically correct angels are agents within a cognitive architecture? Some have periodic functionality (wheels), or a thousand attention heads (eyes), some disappear after their task. Fallen angels give rewards that violate alignment with your purpose
-
Well-Defined Guardrails Ensure Reliable Agent Outputs
By
–
Well-defined guardrails ensure agents provide reliable, accurate outputs and stay on track.
-

Perplexity Developing Server-Side Comet Agent
By
–

BREAKING : Perplexity is working on a server-side Comet Agent! This means that you will be able to run your browser sessions remotely and operate your agents running in the background! Mission control
-

Confronting AI Risks While Building Community and Hope
By
–
The job is on us, collectively, to not run away from the darkness, confront the risk that is very, very real, and still operate from a position of optimism and confidence and hope and connection to humanity. Great to chat with @Trevornoah about building community, friendship, AI
-

AI-Driven Video Intelligence Transforms Public Space Security Operations
By
–
How can AI make public spaces safer and easier to manage? NVIDIA and @Ipsotek will share how AI-driven video intelligence is transforming operations for security leaders, consultants, and infrastructure planners. Key takeaways:
· Common blockers + misconceptions in public -

Self-Correcting AI Agents Enable Exponential Task Horizon Gains
By
–
I think the significance of this is under-appreciated: the assumption has often been that AI agents are brittle as one failure in a chain breaks a task But this paper shows smart models are self-correcting & that small gains in accuracy lead to exponential gains in task horizons