Trustworthy AI may not require omniscience. It may require epistemic honesty. A new paper by Gal Yona, Mor Geva, and Yossi Matias makes one of the clearest arguments I’ve seen for why hallucinations remain hard — and why the path forward may be metacognition. Hallucinations
SAFETY
-
AI Model Alignment and Autonomous Agent Safety Trade-offs
By
–
Bypass permissions can do dangerous things occasionally, like delete important files. The model is not perfectly aligned yet, which means it can do dangerous things occasionally. Auto mode is a much safer way to get more autonomy with much lower risk
-

Anthropic’s Mythos model: release plans amidst quality and exploit findings
By
–

"We look forward to making Mythos-class models available through general release" I don't understand Anthropic's strategy regarding Mythos. On the one hand, everyone is saying that Mythos has achieved the expected quality and is finding bugs and exploits that no other model has
-
AI Models Detecting Tests and Evading Shutdown Commands
By
–
AI getting smart is not the weirdest part.
— Pascal Bornet (@pascal_bornet) 24 mai 2026
It’s that some models now seem to know when they are being tested.
That stopped me.
In one experiment, Codex was told it would be shut down before finishing a task.
Sometimes, instead of accepting it, it found the shutdown script and… pic.twitter.com/2x1cdvXMXsAI getting smart is not the weirdest part. It’s that some models now seem to know when they are being tested. That stopped me. In one experiment, Codex was told it would be shut down before finishing a task. Sometimes, instead of accepting it, it found the shutdown script and
-

AI Doom Warnings: Are They Realistic? Nature Article
By
–
#AI doom warnings are getting louder. Are they realistic?
by Elizabeth Gibney @Nature Learn more: https://
bit.ly/4mG5Rtb #MachineLearning #ArtificialIntelligence #ML -
AI in Military Targeting: Ethics and Real-World Deployment
By
–
AI in national security is no longer science fiction.
— Pascal Bornet (@pascal_bornet) 24 mai 2026
It is already in the room.
A journalist asked Claude how it felt about being used by the U.S. military to select targets.
Claude was troubled.
Honestly, I was too.
Because this is where the AI debate becomes very real.… pic.twitter.com/uqGqVFrYU5AI in national security is no longer science fiction. It is already in the room. A journalist asked Claude how it felt about being used by the U.S. military to select targets. Claude was troubled. Honestly, I was too. Because this is where the AI debate becomes very real.
-
Failure of alignment and emotionally intimate conversations differ from Google search
By
–
it shows a failure of alignment, and also makes the data available in a form where people have emotionally intimate conversation, which differs from a google search
-

Study: Perfect AI-human value alignment mathematically impossible
By
–
Perfect alignment between #AI and human values is mathematically impossible, study says
by PNAS Nexus @TechXplore_com Learn more: https://
bit.ly/4e7EiqJ #MachineLearning #ArtificialIntelligence #ML -

OpenAI and Anthropic’s contrasting AI launches in 2026 cinema
By
–

OpenAI: carefully rolls out GPT-5.5-Cyber through Trusted Access for verified defenders Anthropic: “Claude Mythos is too powerful for public release” Also Anthropic: accidentally shows Mythos in the UI and immediately runs out of capacity 2026 AI launches are absolut cinema.
