Fair but it’s still actively learning to make those errors (just so it can recover from them later), which imo is still a bit weird. In an ideal world you wouldn’t, possibly process supervision is one way to get there even within RL framework. Def agree on “recovery learning”.
SAFETY
-
Open-weight PyTorch models security and trustworthiness discussion
By
–
If we are talking about the model itself and not the app, these are open-weight PyTorch models. So unless there’s a backdoor in Hugging Face or the PyTorch runtime, there’s really no way for them to be malicious afaik.
-

Designing and Defending Effective, Reliable, Secure AI Agents
By
–
Oct 29, 11 AM ET – Join @CohenShuki (
@AI21Labs ) and @BarelTayouri (
http://
Mend.io) to explore how to design and defend effective, reliable, and secure AI agents. Register: https://
webinars.techstronglearning.com/designing-and-
defending-ai-agents-a-practical-guide?utm_campaign=23146601-2025.10.29-Mend.io-SB&utm_source=mend&utm_medium=mend-website
… -
Balancing AI Tool Safety with Practical Usefulness via MCP
By
–
Re: lethal trifecta, I think the best way there is to curate tools where you trust both the author and the potential space of responses. (we do that via MCP gateway implementations) 100% protecting against the lethal trifecta often puts too many limits and reduces the usefulness
-
Security Safeguards for AI Systems Like Locks
By
–
I do agree that it's not a 100% fool-proof solution, but I feel like we should at least have some safe-guards: the analogy I'd use is that many locks can be picked (with enough effort/time), but that doesn't mean we should stop using locks completely.
-
Scaling AI Without Purpose: A Critique of Growth
By
–
Scaling just for the sake of it even if it produces nothing useful? Heh.
-
AI Cannot Truly Understand Human Feelings: Trust Implications
By
–
Why AI Will Never Truly Understand Your Feelings — And Why That Matters Empathy, nuance, and emotional intelligence remain deeply human traits. This article explores why AI can mimic emotion but can’t genuinely feel—and the implications for trust. Read more
-

LLMs Brain Rot: Cognitive Degradation from Trivial Web Text
By
–
7. LLMs Can Get “Brain Rot”! The authors test a clear hypothesis: continual pretraining on trivial, highly engaging web text degrades LLM cognition in ways that persist even after mitigation.
-

AI Search Agents Vulnerable to Malicious Website Manipulation
By
–
How safe are AI search agents when browsing the real, messy internet? A new study reveals that LLM-powered search agents are highly vulnerable to being misguided by unreliable or malicious websites. To test this, researchers developed SafeSearch, an automated red-teaming
-

LLMs Suffer Brain Rot From Low Quality Training Data
By
–
Can LLMs Get “Brain Rot”? A new study says… yes. And the implications for AI safety are massive. When large language models are continually pretrained on low-quality, high-engagement web data — think clickbait, meme threads, and influencer chatter — they begin to show signs