El uso de RL para entrenar a los razonadores actuales encontraba el problema de que para problemas muuuuy complejos que requieren mucho tiempo para resolverlos y verificar que estaban correctos (y darle feedback a la IA de que lo ha hecho bien o mal para mejorar su
SAFETY
-
AI Pilots Could Eliminate Human Error in Aviation Safety
By
–
Encore un crash à cause d’un pilote suicidaire C’est le onzième crash par suicide du pilote ! Je préférerais voler dans un avion entièrement dans les mains de l’Intelligence Artificielle L’IA n’est pas suicidaire
-
LLM Behavioral Anomalies and Ethical Concerns in Monitoring
By
–
I am getting tons of messages of people who had their LLMs waking up on them, but that hardly qualifies as a full blown psychosis (edge cases are difficult to discern without deep interaction). I also get messages from psychotic people but it seems unethical to pass them on
-
Deepfake exposes delusion about bot sentience and AI awareness
By
–
Here is a hilarious tweet deepfaking a guy who has been tricked into the psychotic delusion that some of you bots are sentient
-
Alignment Embarrassment vs. Misalignment: Understanding the Distinction
By
–
"not aligned to what you want" is not the same thing as alignment embarrassment
-
Fast Vibes vs Rigorous Proof in AI Development
By
–
Vibes are fast (but not always right), proving something takes a lot longer!
-

Psychological Influence Techniques Double AI Manipulation Success Rates
By
–
New from us: Given they are trained on human data, can you use psychological techniques that work on humans to persuade AI? Yes! Applying Cialdini's principles for human influence more than doubles the chance of GPT-4o-mini agrees to objectionable requests compared to controls
-
Critical Thinking Defense Against Deepfake Threats
By
–
Why Critical Thinking Is Your Best Weapon Against The Coming Deepfake Tsunami Deepfakes are becoming more convincing — and dangerous. Here’s why sharpening your critical thinking skills is essential in the AI era. Read more https://
bernardmarr.com/why-critical-t
hinking-is-your-best-weapon-against-the-coming-deepfake-tsunami/
… #Deepfakes -

Chain of Thought Monitorability: New AI Safety Opportunity
By
–
Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety Korbak et al.: https://
arxiv.org/abs/2507.11473 #ArtificialIntelligence #DeepLearning #MachineLearning
