New Anthropic research: we find that probing, a simple interpretability technique, can detect when backdoored "sleeper agent" models are about to behave dangerously, after they pretend to be safe in training. Check out our first alignment blog post here: https://
anthropic.com/research/probe
s-catch-sleeper-agents
…
ETHICS
-

Anthropic Research: Probes Detect Backdoored Sleeper Agent Models
By
–
-
AI-Powered Robot Improves Safety for U.S. Recycling Workers
By
–
Thanks to this AI-powered #robot, 1.25 million recycling workers in the U.S. can have a safer working environment.
— Harold Sinnott #MWC26 (@HaroldSinnott) 23 avril 2024
via @gigadgets_ #recycling #Sustainability #robotics #AI #IoT #5G #FutureOfWork @gvalan @Hal_Good
pic.twitter.com/xd0TX4AyvnThanks to this AI-powered #robot, 1.25 million recycling workers in the U.S. can have a safer working environment. via @gigadgets_ #recycling #Sustainability #robotics #AI #IoT #5G #FutureOfWork @gvalan @Hal_Good
-
Dataset Missing Data: Privacy Concerns or Credibility Issues
By
–
dataset didnt have it. not sure if uncollected due to privacy concerns or due to what could happen to the credibility if we knew what people were actually typing in
-
Existential AI Risks: Humanity’s Uncertain Future Ahead
By
–
yeah we’re f’ed. Better enjoy the new nice years left.
-
LLM Safety: Instruction Hierarchy Defense and Team Hiring
By
–
LLMs process text from multiple sources and may face conflicting instructions. We teach our models to follow instructions from the highest priority input, giving better defense against attacks. Our Safety Systems team is hiring:
-
Social Media Authoritarianism and Current World Politics
By
–
Yes, and I wrote more about why this is a particularly bad idea particularly right now here https://
lpeproject.org/blog/social-me
dia-authoritarianism-and-the-world-as-it-is/
… -
Deliberate Terminology Confusion in Technology Industries
By
–
Sometimes I think the terms and abbreviations are deliberately designed to confuse others
-

AI Safety by Design Prioritizes Child Protection Online
By
–
We commit to @thorn and @AllTechIsHuman
's Safety by Design principles to ensure child safety is prioritized in the development and deployment of AI tools. -
Cynefin Framework: Adapting Responses to Contextual Circumstances
By
–
"Using the Cynefin framework is as simple as checking our circumstances before responding to them, questioning our instincts. Is it raining outside or sunny? Is it hailing golf balls? Make sure to adapt your response to what’s happening now.
-

Historical Technology Errors and Intellectual Property Practices Examined
By
–
Apparently new information technologies have been beset by egregious errors and dubious IP practices for at least half a millennium.