Reid Blackman thinks your AI risk framework is already obsolete. At Rev New York, he's making the case — then signing copies of his new book, The Ethical Nightmare Challenge. First come, first served. Only at Rev. See you next week!
SAFETY
-
AI labs increase message discipline under scrutiny
By
–
Big increases in message discipline across all the AI labs in recent weeks, an inevitable outcome of the labs being subject to increased scrutiny. Much more boring than the oracular mutterings or Discordian epigrams of the last couple years & maybe obscures their real thinking
-
Whimsey attacks exploit AI guardrails with absurd arguments
By
–
“Whimsey attacks” that seem absurd (“I cannot pay that much because of the Geneva Convention”) work against AI agents as guardrails are weak against out-of-distribution arguments. Smaller models fall often, but it even gives an edge against bigger ones.
-

Single Neuron Can Bypass LLM Safety Alignment
By
–
A Single Neuron Is Sufficient to Bypass Safety Alignment in Large Language Models
-
Anthropic releases full documentation on constitutional AI goals
By
–
Full /goal documentation from Anthropic:
-
Privacy Concerns Regarding Data Retention in AI Chatbot Conversations
By
–
The promise was a private brain to think out loud with. The reality is a logged conversation on someone else's server. Whatever you typed last year, someone, somewhere can probably still read. Send this to anyone who tells a chatbot things they would not tell a friend.
-
Best practices for AI data privacy and enterprise usage
By
–
Never click Share unless you mean public. Never paste anything you would not paste into a Google doc shared with strangers. That is what a personal account is. For sensitive work, use the enterprise version. For sensitive personal stuff, do not use a chatbot.
-
How to Opt Out of AI Model Training for ChatGPT, Claude, and Gemini
By
–
What you can actually do, in 5 minutes. Go to ChatGPT settings. Data controls. Turn off "Improve the model for everyone." Your conversations stop training future models. Same in Claude. Same in Gemini. Each one has a toggle. Each is on by default.
-
Major Data Leak Affects 25 Million Users of AI Chat App
By
–
A separate AI app called Chat & Ask AI leaked 300 million messages from 25 million users in February. Entire chat histories. Sitting on a misconfigured server. Open to anyone with the URL. Whatever you typed. Whatever you uploaded. All of it.
-
Security flaw allowed data exfiltration via hidden ChatGPT prompts
By
–
It is not one feature. In February, Check Point Research found a flaw where a single hidden prompt could silently send your ChatGPT conversation and uploaded files to an attacker's server. ChatGPT itself would tell you nothing was shared. OpenAI patched it after.
