Rewatching the Jurassic World movie scene and I understand intellectually that it is there because they need a movie to happen but ahhhh every five seconds there is another teachable moment for Ops on Things No One In This Job Should Ever Do.
SAFETY
-
LLM World Models and Failure Case Defense Mechanisms
By
–
Sometimes people say that the LLMs make obvious errors because they don’t have a world model, but one is welcome to one’s guess as to what percentage of population has sufficiently detailed world model to discuss layers of defense against failure cases in a remotely operated…
-
Mechanical Failsafe System: Remote Control and Lock States
By
–
understanding of what the mechanical failsafe which saves the day *actually does*, because it is elsewhere implied that it can be actuated remotely by Ops and that it moved to a locked state versus almost always being in a locked state.
-
Security Seams: Rapid Deployment and Access Control Gaps
By
–
“If you go looking for the seams, what do you find?” Example: canary deploy to fleet-wide rollout in 12 seconds per timestamps, no description of why someone is attempting to access raptor enclosure during a shift change (or any discussion or motivation at all), and mysterious
-

Google DeepMind CEO Warns Society Unprepared for AGI Arrival
By
–
"L'AGI arrive, la société n'est pas prête" Le PDG de Google DeepMind, Demis Hassabis, avertit que l'Intelligence Artificielle Générale (AGI) pourrait arriver dans 5 à 10 ans, mais la société reste non préparée à son impact transformateur. Il plaide pour une collaboration
-
OpenAI Rolls Back Sycophantic GPT-4o Update, Improves Search
By
–
OpenAI a annoncé deux mises à jour majeures cette semaine :
— VISION IA (@vision_ia) 5 mai 2025
— D’abord, une modification du "caractère" de GPT-4o, qui l’a rendu trop flatteur et complaisant… à tel point qu’ils ont finalement fait marche arrière.
— Ensuite, une série d’améliorations pour la recherche dans… pic.twitter.com/XY1PBjT22gOpenAI a annoncé deux mises à jour majeures cette semaine : — D’abord, une modification du "caractère" de GPT-4o, qui l’a rendu trop flatteur et complaisant… à tel point qu’ils ont finalement fait marche arrière.
— Ensuite, une série d’améliorations pour la recherche dans -
Meta LlamaCon: Free API, AI App, and Safety Tools Announced
By
–
Meta a organisé sa toute première conférence LlamaCon dédiée aux développeurs et a fait une avalanche d'annonces, dont :
— VISION IA (@vision_ia) 5 mai 2025
— Un aperçu gratuit de l’API Llama
— Une application "Meta AI" façon ChatGPT avec un fil Discover
— Lama Guard 4 (12B), LlamaFirewall et Prompt Guard pour la… pic.twitter.com/K6KWpaBxqkMeta a organisé sa toute première conférence LlamaCon dédiée aux développeurs et a fait une avalanche d'annonces, dont : — Un aperçu gratuit de l’API Llama
— Une application "Meta AI" façon ChatGPT avec un fil Discover
— Lama Guard 4 (12B), LlamaFirewall et Prompt Guard pour la -

Study Warns AI Control Drops to 52% as AGI Approaches
By
–
Une nouvelle étude sur la "surveillance évolutive" de l'IA révèle des chiffres alarmants : même dans le meilleur des cas, nous ne contrôlons les IA avancées que 52% du temps, avec une efficacité qui diminue en approchant l'AGI. Max Tegmark estime à >90% la probabilité que la
-
Addressing Verbosity and Reliability Feedback for Claude 3.7
By
–
Appreciate the feedback! Cost/verbosity: a few verbosity control workstreams in flight to help manage output length and costs bc I agree 3.7s overeagerness is a little too much at times Reliability: I hear ya on this – reliability issues aren't acceptable. Our team is