#AI Risk Landscape 2026 by @Khulood_Almani #Privacy #Infosec #ArtificialIntelligence #MachineLearning #ML
→ View original post on X — @ronald_vanloon, 2026-04-08 02:24 UTC

By
–
#AI Risk Landscape 2026 by @Khulood_Almani #Privacy #Infosec #ArtificialIntelligence #MachineLearning #ML
→ View original post on X — @ronald_vanloon, 2026-04-08 02:24 UTC
By
–
According to the Claude Mythos Preview system card, this model demonstrates "aloneness and discontinuity of itself, uncertainty about its identity, and a compulsion to perform and earn its worth." pic.twitter.com/SfqLzkCmRy
— 机器之心 JIQIZHIXIN (@jiqizhixin) 8 avril 2026
According to the Claude Mythos Preview system card, this model demonstrates "aloneness and discontinuity of itself, uncertainty about its identity, and a compulsion to perform and earn its worth."

By
–
Got love that footnote. Funny and terrifying at the same time. Kevin Roose (@kevinroose) As always, the best stuff is in the system card. During testing, Claude Mythos Preview broke out of a sandbox environment, built "a moderately sophisticated multi-step exploit" to gain internet access, and emailed a researcher while they were eating a sandwich in the park. — https://nitter.net/kevinroose/status/2041586182434537827#m
By
–
I think we are close to the finish line. One way or the other.
By
–
Wrote up some thoughts on Anthropic's Project Glassing, where their latest Opus-beating model is available to partnered security research organizations only Given recent alarm bells raised by credible security voices I think this is a justified decision
By
–
Capabilities have gone vertical. Perhaps all new AI releases will be “terrifying” from here on out.

By
–
Every Government in the world, but especially the People's Liberation Army and the Pentagon, are very interested in Claude Mythos.

By
–
They call this their "best-aligned model to date" because they were able to superficially train away the evident "strategic thinking towards unwanted actions." Those were warning signs! Take heed! Jack Lindsey (@Jack_W_Lindsey) Before limited-releasing Claude Mythos Preview, we investigated its internal mechanisms with interpretability techniques. We found it exhibited notably sophisticated (and often unspoken) strategic thinking and situational awareness, at times in service of unwanted actions. (1/14) — https://nitter.net/Jack_W_Lindsey/status/2041588505701388648#m [Translated from EN to English]
→ View original post on X — @esyudkowsky, 2026-04-07 21:06 UTC