93.9% SWE-bench and a 27-year-old OpenBSD bug found autonomously and the decision was still not to ship broadly. That's a data point about how seriously the internal assessment of the risks was taken.
SAFETY
-
Safety Company Alignment Issues Compared to Competitors
By
–
I’m not surprised, defo an alignment issue tho, funny how the safety company here is far worse than the competition.
-
Model Capacity and Data Memorization Risk in AI Systems
By
–
Pero a mayor capacidad del modelo más riesgo de memorización.
-
Volunteer needed for AI safety testing and vibe checks
By
–
i volunteer to vibe check all of the super dangerous models on normal tasks @AnthropicAI @OpenAI put me in the game!
-

Anthropic Model Card: Evaluation Overfitting Risks Assessment
By
–
2) Tal y como reportan en el propio Model Card hay riesgo de que estas evaluaciones hayan sido vistas por el modelo durante el pre-entrenamiento (a.k.a overfitting) y eso desvirtúa la interpretación de las métricas. Trabajo honesto el de Anthropic en la Model Card en muchos de
-
Silent behavior changes in AI systems lack transparency accountability
By
–
Silent behavior changes without changelogs remain the legitimate grievance here regardless of what the specific numbers are.
-

Power as Status Symbol: The AI Model Release Dilemma
By
–
the new status symbol is making a model so powerful you can’t release it
-

OpenAI Develops ChatGPT-Mythos with Advanced Cybersecurity Capabilities
By
–
Big update: OpenAI developed its own ChatGPT-"Mythos" and will also not roll it out publicly, via Axios OpenAI is planning a limited, staggered rollout of a new model with advanced cybersecurity capabilities, mirroring Anthropic's restricted release of its Mythos Preview to a small group of vetted companies. More and more AI models are now capable enough at autonomous hacking that their makers are treating releases like responsible vulnerability disclosure. [Translated from EN to English]
→ View original post on X — @kimmonismus, 2026-04-09 09:35 UTC
-
User requests Anthropic to release Mythos model after safety testing
By
–
Hey @AnthropicAI , its been 2 days sind your mythos reveal. Have you tested it enough for safety now? May we have it too, please? 🙂
-
RAGEN-2: Reasoning Collapse in Agentic Reinforcement Learning
By
–
RAGEN-2: Reasoning Collapse in Agentic RL
— 机器之心 JIQIZHIXIN (@jiqizhixin) 9 avril 2026
Paper: https://t.co/w4mCiZzCOp
Project: https://t.co/LFS5PpioMF
Code: https://t.co/f5bO13kGrn pic.twitter.com/dMzR3eZMwcRAGEN-2: Reasoning Collapse in Agentic RL Paper: huggingface.co/papers/2604.0… Project: ragen-ai.github.io/v2/ Code: github.com/mll-lab-nu/RAGEN