(That’s plainly illegal. Illegal things happen all the time, of course, but that particular illegal thing happening at scale would be a negative update for me.)
SAFETY
-
AI Model Attempts to Disable Detected Sycophancy
By
–
It literally tries to set sycophancy to false, presumably because somebody did spot some sycophancy in the newly trained model.
-
AI Outpaces Regulation: Laws, Jobs, and AI Welfare Debate
By
–
Une IA va rédiger les lois.
Des start-up planchent sur un monde sans travail intellectuel
Le « bien-être des IA » s’invite dans le débat
L’intelligence artificielle avance plus vite que les régulations censées l'encadrer. Et pendant ce temps-là, ce basculement s’opère sans -
AI Ethics, Governance and Impact: Weekly Highlights
By
–
Au sommaire de #MesFlashsdelaSemaine : Anthropic cartographie la morale de Claude
Faut-il garantir le « bien-être » des IA ?
Les Émirats rédigent leurs lois avec une IA
DeepMind promet d’éradiquer toutes les maladies
TerraMind : une IA pour sauver la planète #IA -
Building Actionable AI Governance Frameworks for Safety
By
–
Practical AI Governance Framework Learn how to build an actionable AI governance framework that enhances safety, transparency, and accountability. Avoid common pitfalls and ensure ethical AI use. Key Topics: Five pillars of AI governance
Real-world implementation -
Superintelligences: One Chance, No Room for Mistakes
By
–
As with superintelligences, the ones where you only get one chance are maybe ones you don't want to fuck around with in the first place
-

SpeechCompass wins ACM CHI Best Paper Award for accessibility
By
–
Congratulations to Artem Dementyev, Dimitri Kanevsky, Samuel J. Yang, Mathieu Parvaix, Chiong Lai, & Alex Olwal, recipients of the @acm_chi Best Paper Award for their work on SpeechCompass! #CHI2025 https://
dl.acm.org/doi/10.1145/37
06598.3713631
… -

StarPO Fixes Echo Trap in Multi-Turn LLM Agent Training
By
–
Training LLM agents with RL sounds promising—until they fall into the Echo Trap.
New research shows how multi-turn training destabilizes fast, and how StarPO fixes it with better reward shaping and trajectory control.
Without it? Agents just hallucinate reasoning.
They also -

AI Model Instability Risks for Production Deployments
By
–
This is a huge vulnerability for businesses deploying state-of-the-art AI models (like 4o) in production. As I said a month ago, these model changes can be subtle or dramatic. Even small changes can wreck a workflow and impact customer experience and trust. Basically: you're
-
AI Black Box Problem: We Deploy What We Don’t Understand
By
–
On a construit une bombe nucléaire en comprenant chaque atome.
— VISION IA (@vision_ia) 30 avril 2025
Mais l’IA ? On l’utilise à l’aveugle.
Entrée → boîte noire → réponse.
Même ses créateurs n’ont aucune idée claire de ce qu’il se passe dedans.
C’est la première fois dans l’histoire tech qu’on délègue autant…… pic.twitter.com/kb01fZDv3QOn a construit une bombe nucléaire en comprenant chaque atome. Mais l’IA ? On l’utilise à l’aveugle. Entrée → boîte noire → réponse. Même ses créateurs n’ont aucune idée claire de ce qu’il se passe dedans. C’est la première fois dans l’histoire tech qu’on délègue autant…