It isn't about "simple" utility functions or "monomania". The problem is just any sufficiently smart system whose work, on some level, can be viewed as matching up outputs and results, and learning.
ETHICS
-
Learning Reality and Selecting Outputs for Desired Outcomes
By
–
The problem that 'learn how reality works' and 'select outputs which, when they interact with reality, lead to X happening' is a simple great way of doing Y for a lot of possible Y. For example, with humans, Y is inclusive genetic fitness and X is all the stuff that humans want.
-
KQV matrices: wrong level of abstraction for AGI safety concerns
By
–
If you want something about kqv matrices, you're asking for an explanation on the wrong level of abstraction; if the problem was specific to kqv matrices we'd advocate "stop using transformer layers" not "shut down AGI research".
-
Utility Functions and Existential Risk from Instrumental Convergence
By
–
https://
arbital.com/p/instrumental
_convergence/
… but I'm not sure what that buys you if "pick any simple measure on utility functions, preimage them through a reasonable environmental model onto actions, most utility functions kill humanity as a side effect" doesn't already do it. -
Bing Sydney’s Unexplained Threat Behavior Remains Mystery
By
–
Nope! To this day, as far as I know, nobody knows which particular numbers inside Bing Sydney led her to try to threaten a human with reporting him to the police. They poked her until she stopped doing that, but can't read her thoughts any more than we can look at a frozen
-
Neural Networks: Billions of Inscrutable Numbers We Cannot Understand
By
–
Nope! Neural nets are built by repeatedly poking a table of hundreds of billions of inscrutable numbers until they start doing what the builders want. We understand the thing that does the poking, but not what the hundreds of billions of numbers mean.
-
Preference Functions and Alignment: Optimizing for Human Flourishing
By
–
If you don't see why most preference functions that don't specifically have an attainable optimum around people living happily ever after, do something else instead of that, I'm not sure what particular inscrutable properties of a kqv layer are going to be persuasive?
-
Instrumental Convergence Denial Prevents Advanced AI Safety Discussion
By
–
Yann is currently at the stage of denying instrumental convergence, so there's no point in bringing in anything more complicated from List of Lethalities.
-
Debating AI Safety: The Unbridgeable Gap in Risk Perception
By
–
If someone performs that they can't see any difference between building a poorly understood superhuman intelligence, and a 20th-century newspaper article inveighing against coffee, they are beyond the reach of debate. I can only go to the general public and say, "This is their
-
Contractual Obligations and Copyright in AI Development
By
–
Contractual obligation doesn't require Copyright.