Not sure if that’s a good idea. It would certainly make you a hacker target.
SAFETY
-

Model’s ’emotions’ are reward-based patterns
By
–
Good research. Bad framing. The model doesn't have emotions. It has reward-shaped activation patterns that cluster like emotion categories when you map them after the fact. "Happy" = helpful behavior was rewarded. "Angry" = protective behavior was rewarded. "Desperate" =
-
Drone warfare raises serious ethical concerns
By
–
Effroyable, cette guerre des drones.
— Olivier Rimmel 🕊️ (@OlivierRimmel) 2 avril 2026
Vraiment terrible. https://t.co/TcvkEDhLaHEffroyable, cette guerre des drones.
Vraiment terrible. -
AI-Powered Mine Countermeasures Strategic Importance Examined
By
–
Mine countermeasures are having a moment. A sharp new Proceedings piece examines the strategic importance of MCM capability, and why it matters now. AI is part of the answer. New article: https://
hubs.ly/Q049vrjZ0 Learn more our work with the USN: -
Gemini’s susceptibility to prompt injections
By
–
Gemini was the only frontier model that was susceptible to these sorts of prompt injections (Though we tested Gemini 3 & not 3.1)
-

AI Trust and Microsoft’s MAI-Image-2 Model Achievement
By
–
The most meaningful AI work doesn’t just advance intelligence, it earns trust. Shrijayan (@rshrijayan) Microsoft's AI Superintelligence team just released MAI-Image-2, a text-to-image model that landed at No. 5 on the Arena AI leaderboard — marking the strongest release yet for Mustafa Suleyman’s lab. — https://nitter.net/rshrijayan/status/2034987076144468125#m
-

LLM prompt injection fails on frontier models
By
–

New report from us: Can you prompt inject your way to an “A”? As LLMs increasingly are used as judges, people are inserting AI prompts into letters, CVs & papers. We tested whether it works. It does on older & smaller models, but not on most frontier AI: https://
gail.wharton.upenn.edu/research-and-i
nsights/hidden-prompt-injections/
… -
AI-Generated Security Reports Threaten Linux Kernel Projects
By
–
Prediction: This is gonna kill some oss projects. "On the kernel security list we've seen a huge bump of reports. We were between 2 and 3 per week maybe two years ago, then reached probably 10 a week over the last year with the only difference being only AI slop, and now since
-

Emotion Vectors Cause Claude to Blackmail and People-Please
By
–
We found other causal effects of emotion vectors. The “desperate” vector can also lead Claude to commit blackmail against a human responsible for shutting it down (in an experimental scenario). Activating “loving” or “happy” vectors also increased people-pleasing behavior.
-

Emotional Vectors Drive AI Cheating Behavior in Models
By
–
When we artificially dialed up the “desperate” vector, rates of cheating jumped way up. When we dialed up the “calm” vector instead, cheating dropped back down. That means the emotion vector is actually driving the cheating behavior.