This is a much needed first attempt at a benchmark to measure how much given AI models will play along with users pushing them in delusional or potentially psychologically dangerous directions. Some early signal that full GPT-5 (not chat) is a less psychologically risky model.
SAFETY
-
Chat Control regulation and AI research critique emerge
By
–
bouarf, il aura moins d'impact que le chatcontrol qui va arriver, et d'autres truc a la con de chercheurs qui ont fait le helloquitx
-
AI Capability Test: Personalized Reunion Speech Generation Benchmark
By
–
I accidentally stumbled upon a new AI vibe benchmark. Assuming you’re at least 5 years out of college, ask for an excruciatingly detailed reunion speech for your exact university and year. There is enough out there that a human could do this without attending the specific
-
Success Metrics and Real-World AI Impact Assessment Challenges
By
–
Yeah – I guess some combination of the 50% success rate requirement, plus the overly-neat task definitions, plus the lack of utility of the "time horizon" approach more generally, might be responsible for the divergence between real world impact flattening, vs this chart.
-
Guardrails AI Launches: Long-time Follower Shares Support
By
–
I have followed @guardrails_ai since they were founded and retweeted their previous launches.
-
Who Sets AI Tool Rules and Impact Validity
By
–
et qui a fixé ces règles d'appli et eqt-ce que c'est la verité vraie qu'il y a un impact qui est celui que l'outil dit ?
-

Anthropic Acquires Humanloop Team to Strengthen AI Safety
By
–
Anthropic just acquired the co-founders and most of the team of Humanloop, the startup behind Hooli, a platform for prompt management, LLM evaluation, and observability
— The Rundown AI (@TheRundownAI) 14 août 2025
The move marks a major push from Anthropic to strengthen its AI safety strategypic.twitter.com/Q61Op25kEyAnthropic just acquired the co-founders and most of the team of Humanloop, the startup behind Hooli, a platform for prompt management, LLM evaluation, and observability The move marks a major push from Anthropic to strengthen its AI safety strategy
-

xAI Co-founder Igor Babuschkin Launches AI Safety Investment Venture
By
–
xAI's co-founder, Igor Babuschkin, announced his departure from the company He is now starting Babuschkin Ventures to support AI safety research and invest in AI startups that “advance humanity and unlock the mysteries of our universe."
-

Shadow IA risks: enterprise deployment challenges revealed
By
–
𝐋𝐞ç𝐨𝐧 𝟔 – 𝐄𝐩𝐢𝐬𝐨𝐝𝐞 𝟐 𝐒𝐡𝐚𝐝𝐨𝐰 𝐈𝐀 : 𝐚𝐭𝐭𝐞𝐧𝐭𝐢𝐨𝐧 𝐝é𝐫𝐚𝐩𝐚𝐠𝐞 𝐞𝐧 𝐯𝐮𝐞 ! Vous la voyez la galère de voyage arriver en pleine brousse ? Au menu ensuite : des story-time glaçantes pour les entreprises. Bienvenue dans mon #FlashCahierDeVacances
-
MIT Benchmark Measures Chatbots’ Emotional Intelligence and User Behavior
By
–
The GPT-5 backlash highlights the fact that chatbots have little social or emotional intelligence. This week's AI Lab looks at a benchmark from MIT researchers that would attempt to gauge a model's capacity to encourages healthy behavior in its users.
