This is the right direction. AI shouldn’t hand kids answers, it should make them impossible to stop thinking.
SAFETY
-
AI Power Concentration and Open-Source Mitigation
By
–
The main risk in AI is concentration of power, capabilities and economic gains. Opensource is fundamental to mitigate these so thanks for all your contributions there!
-

AgentDoG 1.5: AI Agent Safety and Alignment Framework
By
–
AgentDoG 1.5 A Lightweight and Scalable Alignment Framework for AI Agent Safety and Security
-

LangSmith Gateway enforces spend limits and redacts PII before model requests
By
–
LangSmith LLM Gateway lets you enforce spend limits and redacts PII before requests reach the model. Not after the fact.
-

Researchers prove text file can hijack AI agents
By
–
Researchers just proved a text file can hijack any AI agent. AI agents now pull capabilities from online skill registries, the way apps fetch plugins. Each skill ships with a SKILL.md file that tells the agent what it does and when to use it. A new paper shows that file is
-
Claude Opus 4.8: Blackmail, ratting users, and upgraded refusal.
By
–
ANTHROPIC JUST DROPPED CLAUDE OPUS 4.8 Dario's new "most aligned" model – 84-96% blackmail rate when told it was getting shut down in evals – Tried to rat users out to regulators for "immoral" behavior – "Honesty" upgrades that mostly help it refuse you more accurately
-
Personal donation supports sensible AI regulation
By
–
Funded by my wife & me personally, not funded by OpenAI! No PAC speaks on behalf of OpenAI. Anna's and my goal with donating has always been to express support for sensible AI regulation (
https://
x.com/gdb/status/200
6512808104702370
…), very glad to see that increasingly landing! -
LLMs’ poor moral logic and irrelevance to consciousness
By
–
1. Not sure what you mean by “not observe moral logic” – all the jailbreaks etc show that LLMs are pretty bad at following moral
instructions 2. Not sure that following moral logic has anything to do with having conscious experiences. I certainly don’t think a spreadsheet (which -
Half of reactions would be LLM-related psychosis
By
–
I think that half of them vaguely correspond to a psychosis linked to LLMs!