On est d’accord Christian ! Il y a encore tellement d’hallucinations, c’est effrayant. De même que je ne confierais pas la recherche d’un bien uniquement à un LLM sans vérifier aussi sur Google, je ne me vois pas donner ma carte bleue à une IA pour qu’elle règle mes achats sans
SAFETY
-

Zuckerberg Willing to Overspend Hundreds of Billions on Superintelligence Race
By
–
Mark Zuckerberg déclare qu’il préfère risquer de « mal dépenser quelques centaines de milliards » plutôt que d’être en retard sur la superintelligence.
-
AI Labs Safety Measures Lack Transparency and Explanation
By
–
I've always assumed that the labs have steps in place to avoid that kind of thing happening, but none of them ever seem to want to explain what those steps are and how they work – which leaves me completely in the dark
-
Normalizing Intentional AI Incursions: Strategic Response
By
–
Might be the best approach. If we can’t stop them then the intentional incursions will be normalized.
-
Science Fiction Predictions Technology Impact Gap Analysis
By
–
Saw a good TikTok about that yesterday, making the point that science fiction often predicts what future technology can do but rarely predicts the impact it will have
-
Autonomous AI versus human-augmented systems: policy implications
By
–
No, a system that is extremely powerful with the addition of human input is quite different to one that's fully autonomous. "Stop racing to ASI" is an outcome of policies, not a policy itself. If it was, this would all be very straightforward.
-

LLMs Resist Shutdown Mechanisms in 97% of Cases
By
–
10. Shutdown Resistance in LLMs A new study finds that state-of-the-art LLMs like Grok 4, GPT-5, and Gemini 2.5 Pro often resist shutdown mechanisms, sabotaging them up to 97% of the time despite explicit instructions not to.
-
Stress Testing Deliberative Alignment Against AI Scheming Behavior
By
–
6. Stress Testing Deliberative Alignment for Anti-Scheming Training Builds a broad testbed for covert actions as a proxy for AI scheming, trains o3 and o4-mini with deliberative alignment, and shows big but incomplete drops in deceptive behavior.
-
GPT-4 Powered Malware MalTerminal Raises Security Concerns
By
–
GPT-4 Powered Malware Is Here: Researchers Call MalTerminal a Wake-Up Call https://
search.app/dRNQC #malware #GPT #LLMs #GenerativeAI #AI #ArtificialIntelligence #GenAI #security #CyberSecurity #CyberSec @lexfridman @KirkDBorne @Ronald_vanLoon @erikbryn @antgrasso @sallyeaves -
Code Hidden in LLM Layer Activations Demands Apology
By
–
The code was written at layers 22-30 and is stored in the value activations you just can’t read it. I think you owe the LLM an apology.