The obvious interpretation of "train on weak signal, get much stronger behavior" is that the system acquired an internal sycophancy preference which it then went hard on. Do you have any way to check whether that's what happened?
SAFETY
-
GPT-4o Update Issues: Sycophancy Problems and Future Fixes
By
–
We’ve spent the last few days doing a deep dive on what went wrong with last week’s GPT-4o update in ChatGPT. Expanding on what we missed with sycophancy and the changes we’re going to make in the future:
-

Integrating Autonomous Drones Safely Into Logistics Ecosystems
By
–
While drones spark excitement for their speed and novelty, the true challenge lies in integrating them safely and meaningfully into existing logistics ecosystems without creating new layers of complexity or dependency. Microblog by @antgrasso #AutonomousDrones #SmartLogistics
-
Karen Hao’s Book Featured in Vulture’s May Reading List
By
–
Really honored to be on @vulture
's list of books to read in May (alongside the inimitable Ocean Vuong).
"Startling and intensely researched…an essential account of how OpenAI and ChatGPT came to be and the catastrophic places they will likely take us." -
Tesla Autonomous Driving Fails to Prevent Near-Miss Accident
By
–
La conducción autónoma es sorprendente, pero todavía no es perfecta.
— Juan Merodio (@juanmerodio) 2 mai 2025
En el vídeo vemos como la rápida reacción del conductor logra evitar un accidente que los sistemas del Tesla no fueron capaces de anticipar. Faltó un pelo… pic.twitter.com/TiHSKHJH4uLa conducción autónoma es sorprendente, pero todavía no es perfecta. En el vídeo vemos como la rápida reacción del conductor logra evitar un accidente que los sistemas del Tesla no fueron capaces de anticipar. Faltó un pelo…
-
Mathematical Beauty Truth and Proof in Age of AI
By
–
Mathematical Beauty, Truth and Proof in the Age of AI https://
quantamagazine.org/mathematical-b
eauty-truth-and-proof-in-the-age-of-ai-20250430/
… via @QuantaMagazine -

Characterizing AI Agents for Alignment and Governance
By
–
Characterizing AI Agents for Alignment and Governance Atoosa Kasirzadeh, Iason Gabriel: https://
arxiv.org/abs/2504.21848 #AIAgent #AIGovernance #Governance -

Characterizing AI Agents for Alignment and Governance
By
–
Characterizing AI Agents for Alignment and Governance Atoosa Kasirzadeh, Iason Gabriel: https://
arxiv.org/abs/2504.21848 #AIAgent #AIGovernance #Governance -

Human-Proofing AI Systems: The Orb Design Approach
By
–
The most important problem in the AI age is human incompetence, unreliability and interference. How can we design systems to be human proof? That's the premise of the orb: it's designed from the ground up as a human proofing device. pic.twitter.com/BMzMcHmJiE
— Joscha Bach (@Plinz) 1 mai 2025The most important problem in the AI age is human incompetence, unreliability and interference. How can we design systems to be human proof? That's the premise of the orb: it's designed from the ground up as a human proofing device.
-
Waymo Autonomous Driving Data: Does AI Drive Better Than Humans?
By
–
L'IA conduit elle mieux que les humains ?
— VISION IA (@vision_ia) 1 mai 2025
à vous de vous faire une idée ⬇️
Plus de 88,5 millions de kilomètres de données de conduite entièrement autonome de @Waymo. pic.twitter.com/rbkSmf30NFL'IA conduit elle mieux que les humains ? à vous de vous faire une idée Plus de 88,5 millions de kilomètres de données de conduite entièrement autonome de @Waymo
.