Simple adversarial attacks in the wild, in war
SAFETY
-
Randomness in AI evaluation makes reproducibility challenging
By
–
That is surprising, but there is a lot of randomness. If you try again from scratch, you may get better (or worse) results. It makes evaluation really hard.
-
GPT-5 Safety Training: Safe Completions Approach
By
–
For GPT-5, we introduced a new form of safety-training — safe completions — which teaches the model to give the most helpful answer where possible while still staying within safety boundaries. If a request can’t be met safely, ChatGPT may partially respond or give a brief
-
Trust Issues with AI Model Updates and Platform Reliability
By
–
I still have trust issue with them over the way they update their models – they need to earn my trust back before I take them seriously as a platform to build my own stuff on top of
-

GPT-5 Launch Reveals Critical Numerical Reasoning Failures
By
–
The GPT 5 launch included a chart showing 52.8 as a bigger number than 69.1, which in turn is shown as the same magnitude as 30.8. Not quite ASI…
-
Framing AI Products: Avoiding Dangerous Perception Messaging
By
–
it helps to not frame your product as a gigantic laser weapon that will blow the planet up
-
Armed Amphibious Robot Dogs Signal Alarming AI Robotics Advance
By
–
La course vers Skynet continue : des chiens amphibies armés tout droit sortis de Black Mirror sont là. pic.twitter.com/XSkjEzgjsH
— VISION IA (@vision_ia) 7 août 2025La course vers Skynet continue : des chiens amphibies armés tout droit sortis de Black Mirror sont là.
-
LLM-Induced Psychosis: Mental Health and AI Safety Implications
By
–
LLM induced psychosis is just regular latent psychosis while being exposed to an environment that is not actively providing stability. GenZ stare is just regular teenage stare while being exposed to an environment that actively discourages adult behavior


