It doesn’t seem like many people work on capability and safety at the same time. I kind of think most people understand they are not the difference maker on whether the consequences turn out good or bad. If they genuinely think it’s the latter, they stop working on capability.
ETHICS
-

AI can design viruses, toxins, and bioweapons: concern level?
By
–
#AI can design viruses, toxins and other bioweapons. How worried should we be?
by Ewen Callaway @Nature Learn more: https://
bit.ly/4eLZseg #ArtificialIntelligence #Innovation #EmergingTech #Biotech -

Beneficial RL data improves AI alignment across tasks
By
–
There are papers that show training AI on "evil" data results in general misalignment, so it is nice to know the opposite is true and that beneficial RL data in one field leads to more aligned models across a range of tasks.
-
Open Weights: Control Over Capability Even When Trailing
By
–
Open weights matter even when they trail the frontier exactly because of this, you're buying control, not just capability. weights matter even when they trail the frontier exactly because of this, you're buying control, not just capability.
-
Ineffective a posteriori API guardrails for cutting-edge models
By
–
Let's face the truth: a posteriori API guardrails are not the appropriate safety tool for cutting-edge models. They do not eliminate dangerous capabilities. They simply hide them behind a fragile interface that can be easily
-
Modèles plus fiables et alignés pour l’IA
By
–
This is an early step toward more robustly beneficial and aligned models: training models to carry beneficial traits into new situations, so as AI becomes more capable, it also becomes more reliable, transparent, and helpful for people.
-

OpenAI tests alignment persistence: model resists harmful prompts, stays helpful
By
–
We also tested whether alignment persisted under pressure. The model was harder to steer toward harmful behavior with adversarial prompts, while remaining responsive to helpful instructions. We saw preliminary evidence of greater resistance to harmful fine-tuning.
-

Cross-domain transfer improves model behavior beyond health conversations
By
–
The most interesting test was cross-domain transfer. When beneficial behavior training was limited to health conversations, the model still improved on non-health evaluations of misalignment, deception, and reward hacking—even though those tasks looked very different from the
-

OpenAI trains models with RL to reinforce beneficial traits across 12 domains
By
–
We trained models with reinforcement learning on realistic conversations to reinforce beneficial traits like truthfulness, humility under uncertainty, openness to correction, fairness, and concern for human welfare, across 12 domains, including health, science, and education.

