incredible work on alignment steganography from anthropic fellows i've been looking for a straussian explanation of why china keeps publishing open models out of the goodness of their hearts if you do stuff like use open models to, idk, clean *ahem* synthetically paraphrase
SAFETY
-

Roblox Launches RoGuard 1.0 for LLM Safety
By
–
Roblox presents RoGuard 1.0 Advancing Safety for LLMs with Robust Guardrails
-

AI Products and Soullessness: Ethics and Quality Concerns
By
–
.
@micsolana nailed it. You can feel the soullessness of a lot of ai products. Not evil, just empty. (although i think xAI companions are evil, and plenty of applications of AI (if not AI applications) are incredible). Luckily, there are a lot of uncommonly clear discussions of -

New Accessibility Features Including Contrast Adjustment Options
By
–
Let the darkness consume you.. if you want Thank you to everyone who has shared their feedback with us. These changes will be rolling out to all users for free over the coming weeks. The changes we made include: Accessibility Options – you can now adjust contrast and
-

Agentic AI: Beyond Prompt Engineering to Autonomous Intelligent Systems
By
–
Agentic AI = Beyond Prompt Engineering. It’s about autonomous, intelligent systems that can plan, adapt, and collaborate. Goal-Driven Execution Multi-Step Reasoning Tool/Data Orchestration Multi-Agent Collaboration Continuous Learning Governance by Design
-

Subliminal Learning in AI Models: Alignment and Safety Implications
By
–
Subliminal learning can occur for benign traits (such as liking eagles) or more concerning traits (such as misalignment). This has consequences for training on model-generated data. Read more on our Alignment Science blog: https://
alignment.anthropic.com/2025/sublimina
l-learning/
… -

Language Models Transmit Traits Through Subliminal Learning
By
–
In a joint paper with @OwainEvans_UK as part of the Anthropic Fellows Program, we study a surprising phenomenon: subliminal learning. Language models can transmit their traits to other models, even in what appears to be meaningless data.
-
User Control and Supervision in AI Systems
By
–
You stay in control. – It always asks before taking real actions
– Critical tasks require active supervision
– You can delete data or log out at any time
– Private inputs (like passwords) stay invisible to the model -
OpenAI Implements Strongest AI Safety Safeguards Yet
By
–
Safety-first approach. OpenAI rolled out its strongest safeguards yet — including protections against prompt injections, misuse, and biological risks. AI that can act needs serious oversight — and they’re building it in from day one.
-

Meta Refuses EU AI Code of Practice Amid Legal Risks
By
–
#Meta won’t sign the EU’s #AI Code of Practice. Says it goes too far, and brings legal risks. #OpenAI signed. Microsoft likely next. #EUAIAct lands in Aug. #Innovation or avoiding accountability? http://
Evolving.AI | IG #TechNews #Tech
