Right, the old "instrumental objective" story.
You can have low-level guardrails against bad effects of instrumental sub-goals.
The question here is not "can you come up with a way that this could go wrong?", but rather "is there a way to do it right?" It's like turbojet design.
ETHICS
-
Instrumental Objectives and Low-Level Guardrails in AI Design
By
–
-
Democratic Process Establishes Good Bad Standards
By
–
We have a democratic process to establish what constitutes "good" and "bad".
It's called lawmaking. -
Regulating Bad AI Through Law Enforcement and Legal Frameworks
By
–
No.
It's like every society: you stop the bad guys AI with laws and law enforcement, i.e. with a well-equiped and well-trained AI police. The American trope refers to people taking the law into their own hands.
Bad idea. -
Strategic Lobbying Behind AI Power: Beyond Emotional Reactions
By
–
To portray the powers at play as reactive and emotion-driven is misleading, as the lobbying force behind them is IMHO strategic and calculating — almost devoid of emotion to the point of psychopathy.
-
AGI Power: Beyond Greed to Governance Control
By
–
OK, challenge accepted! I don't think greed (as pictured) is the main driver for people who stand to benefit from AGI. As a class, they have all the money they want. What they want is the next form of power: governments are decaying, and they need to control what's next…
-
Alignment Concerns with Internal Health Reward System Design
By
–
The first comment (kudos for open review) links to a post that says some of this at greater length, but to repeat my own reaction: "There's nothing in there about alignment. The proposed motivational system is internal-system-health reward with nothing about caring for humans."
-
Two Distinct Metrics for Evaluating AI: Philosophy and Practice
By
–
Great question, but there are really two distinct metrics: one is the more philosophical one: “does AI ‘truly’ reason and think?” Another is a practical one: “what are the economic and societal effects of AI?” Both are important, but they are not the same.
-
Zero Accidents: When Autonomous Technology Becomes Trustworthy
By
–
If the technology becomes so robust and reliable that, over a long time period there is essentially zero accidents, I’d say this is something that could be considered #TeamHuman