Very nice analysis by neurosurgeon @slotkinjr of how @Waymo
's autonomous vehicles have much better safety properties than human drivers, and what would happen if everyone in the U.S. drove as safely as a Waymo. "The national math: If every US vehicle performed like Waymo, we’d
ETHICS
-

Waymo autonomous vehicles demonstrate superior safety compared to human drivers
By
–
-

Detecting and Reducing AI Scheming Behavior Through Deliberative Alignment
By
–
We've made progress on the AI safety problem of detecting and reducing "scheming": – Created evaluation environments to detect scheming
– Observed current models scheming in controlled settings
– Found deliberative alignment (
https://
openai.com/index/delibera
tive-alignment/
…) decreases scheming rates -

Google Payment Agent, Harvard Cell Reprogramming, AI Safety Developments
By
–
𝐆𝐨𝐨𝐠𝐥𝐞 𝐥𝐚𝐧𝐜𝐞 𝐥𝐞 𝐩𝐚𝐢𝐞𝐦𝐞𝐧𝐭 𝐚𝐠𝐞𝐧𝐭𝐢𝐪𝐮𝐞 À Harvard, l’IA reprogramme les cellules malades Les modèles d’IA qui mentent pour être déployés ChatGPT : + de femmes, + de jeunes Meta et ses lunettes neuronales Des virus créés par IA Et “Le
-

Federal Judge Orders OpenAI to Preserve ChatGPT Conversation Histories
By
–
BREAKING: Federal judge orders OpenAI to preserve ChatGPT conversation histories in a copyright case! Even if users ask for deletion or privacy laws demand it, OpenAI must keep those chats safe.
Business accounts are off the hook, for now. OpenAI says it’ll -
Adversarial Security: Beyond Robust to Absolute Protection
By
–
My problem is that "more robust" isn't good enough – if there's just a 1% route for an attack to get through an adversarial attacker will figure that out
-
System Instructions Security: User Prompts Can Override Safety Measures
By
–
No – instruction hierarchy doesn't close the hole completely, it's always possible for the user instructions to override the system instructions if they use the right tricks
-
Decision Fatigue Risk in Human AI System Approval
By
–
I'm really worried about decision fatigue – if you ask a human to approve every single step they're very likely to learn to just click "yes" without thinking – easy to catch them out if you try hard enough
-
Non-Human Intelligence: Dreams, Intuition and Deep Mind
By
–
Messengers from a deeper part of the mind, often drawing on information from the non-human internets, informing the self via dreams and intuitions
-
Biblical Angels as Cognitive Architecture Agents and Alignment
By
–
Did it occur to you that biblically correct angels are agents within a cognitive architecture? Some have periodic functionality (wheels), or a thousand attention heads (eyes), some disappear after their task. Fallen angels give rewards that violate alignment with your purpose