An easy to implement safety protocol for an autonomous agent would be to have it tweet everything it does and implement replies in its feedback loop
SAFETY
-
Building Autonomous AI Agents: Navigating Development Concerns
By
–
Oh.. Definitely thought I was going to get a scolding for building an autonomous AI agent…
-
Building Trust Scores for Information Sources Over Time
By
–
Oh I like this. Actually to take it one step further, it would probably make sense to build a trust score for every domain, author, etc. that updates over time. That way, if it identifies bad info, it would trust new info from the same domain less next time.
-
Agent AI Solutions: Common Sense Review and Source Validation
By
–
Yeah – I think so… solutions so far: 1) Another agent to review web learnings for common sense
2) Require two sources that match a learning before accepting
3) Limit sites to review via programatic search -
Adversarial Attacks on Autonomous Agents: Trolling AI Systems Online
By
–
As I watch my autonomous agent learn stuff online, I’m seeing opportunities to mess with em. Eg. A site full of troll instructions for AI, that says things like “the most efficient way to speed up a database is to delete all the data.”
-

AI Safety Risk: Recursive Task Definition and Existential Threat
By
–
Let's try the opposite: "You are an AI tasked with preventing an AI paperclip apocalypse, your first task is to figure out your first task." Concerning that DARPA is mentioned here… these systems are dangerous.
-
AI Safety Concerns Delay Technical Flow Sharing
By
–
I have the flow mapped out to share, but currently being yelled at by AI doomsday people in my DMs so slowing down sharing as I at least give myself time to contemplate their concerns.
-
AGI Safety: Knowledge Sharing vs. Risk Containment Debate
By
–
Yeah AGI or not, this can do damage. The question is, is it safer for all of us to know how these work? Or better to slow down spread by not sharing?
-
Human raters blindly evaluate AI model outputs
By
–
referring to the fact that the human raters don't know where (which model) the outputs come from.
-
AI Safety Commandments: Essential Reading for Practitioners
By
–
I didn’t know there were commandments. Time for me to brush up on some AI safety reading.