If you can’t stop small teams from using your API for distillation, then you’re definitely not stopping criminals, biohackers, or adversarial states from using AI through it. Those actors are much more sophisticated. The whole "APIs are safer" is a hoax!
SAFETY
-

Unpredictable AGI may resist control, diverse AI safer
By
–
Unpredictable AGI may resist full control, making diverse #AI safer
by Gaby Clark @TechXplore_com Learn more: https://
bit.ly/48empCF #ArtificialIntelligence #MachineLearning #ML -

AI Models Automate Self-Jailbreaking Process Advancement
By
–
I mentioned this project in this week's AI Lab newsletter. It suggests a kind of flip side to the idea of self-improvement in AI models. As models become more powerful, it shows they can automate the process of figuring out how to jailbreak themselves and other models remarkably
-
GPT-5.5 Iterative Deployment Strategy for AI Safety
By
–
1. We believe in iterative deployment; although GPT-5.5 is already a smart model, we expect rapid improvements. Iterative deployment is a big part of our safety strategy; we believe the world will be best equipped to win at the team sport of AI resilience this way. 2. We believe
-
Complex Investigation Reveals Root Causes and System Confounders
By
–
We take these reports incredibly seriously. In my time on the team, this has probably been the most complex investigation we’ve had. The root causes were not obvious, and there were many confounders.
-
Creator of the Term Prompt Injection Shares Its Origin
By
–
I named prompt injection after SQL injection a few years ago
-
Internal recursive self-improvement ASI development approach revealed
By
–
We keep the open-ended recursive self-improvement version internally, for now. #ASIFirst
-
Fraudulent Author Names Discovered in Academic Paper
By
–
all the author names were fraudulent; they had nothing to do with the paper
-

MIT Improves Reasoning Model Confidence Calibration Through RL Training
By
–
How do top reasoning models become overconfident? MIT found that RL rewards correct answers w/o considering how sure the model is. By training them to estimate their confidence about each answer, the team boosted uncertainty estimates w/o hurting accuracy:
