https://
arbital.com/p/instrumental
_convergence/
… but I'm not sure what that buys you if "pick any simple measure on utility functions, preimage them through a reasonable environmental model onto actions, most utility functions kill humanity as a side effect" doesn't already do it.
@esyudkowsky
-
Utility Functions and Existential Risk from Instrumental Convergence
By
–
-
Bing Sydney’s Unexplained Threat Behavior Remains Mystery
By
–
Nope! To this day, as far as I know, nobody knows which particular numbers inside Bing Sydney led her to try to threaten a human with reporting him to the police. They poked her until she stopped doing that, but can't read her thoughts any more than we can look at a frozen
-
Neural Networks: Billions of Inscrutable Numbers We Cannot Understand
By
–
Nope! Neural nets are built by repeatedly poking a table of hundreds of billions of inscrutable numbers until they start doing what the builders want. We understand the thing that does the poking, but not what the hundreds of billions of numbers mean.
-
Preference Functions and Alignment: Optimizing for Human Flourishing
By
–
If you don't see why most preference functions that don't specifically have an attainable optimum around people living happily ever after, do something else instead of that, I'm not sure what particular inscrutable properties of a kqv layer are going to be persuasive?
-
Instrumental Convergence Denial Prevents Advanced AI Safety Discussion
By
–
Yann is currently at the stage of denying instrumental convergence, so there's no point in bringing in anything more complicated from List of Lethalities.
-
Debating AI Safety: The Unbridgeable Gap in Risk Perception
By
–
If someone performs that they can't see any difference between building a poorly understood superhuman intelligence, and a 20th-century newspaper article inveighing against coffee, they are beyond the reach of debate. I can only go to the general public and say, "This is their
-
Energy costs and biological systems: beyond genetic fitness optimization
By
–
Biological systems have energy costs per-computation and this didn't "eliminate" all goals other than pure desire for genetic fitness, because the thousand shards of desire weren't extraneous, they were the whole implementation. But even leaving that aside – how much more
-
Text prediction machinery has cognitive quirks, not ideal simulation
By
–
You should care about weird machinery underlying text predictions that has its own cognitive quirks, not imagine it as an ideal simulacrum plus noise.
-
Outer Optimizers and Inner Optimizers: Beyond Naive Reward Function Pursuit
By
–
From my perspective, the point of raising the example of natural selection is that it debunks the naive belief that if in general an outer optimizer trains on reward function, it gets an inner optimizer that pursues that reward function OOD. Saying "But SGD is first-order and
-
Eliezer Yudkowsky Disputes Misinterpretation of His AI Theory
By
–
I straightforwardly deny that my theory says that GPT gets worse at generalization with more parameters; she doesn't understand where or when my story about an AI deciding to kill us has the murder enter into it. Given that I'm unimpressed by her claim above, what should I think