What do you think of expanding level 4 more and more the way Waymo is doing in the Bay Area?
SAFETY
-
Natural Selection’s Simple Rules Produce Unexpectedly Complex Human Outcomes
By
–
Natural selection is simple. A complicated story then happens, and human beings come out the other end in a way that doesn't match up to the simple stuff natural selection optimized for. I could indeed make this story more complicated; this doesn't help the safety-is-easy case.
-
Complexity in AI Alignment: Why Precision Matters for Safety
By
–
From my perspective, the more complicated you say these relationships are, the worse for anybody trying to build them out in some precise way that doesn't killeveryone. Rockets are complicated and that doesn't make them easier to notexplode.
-
Utility Function Terminology and Human Evolutionary Misalignment
By
–
I think that "utility function" in retrospect is a mathematical word of power that I should not have expected lay computer scientists to understand, so let's drop that. With humanity, the outer optimization criterion was inclusive fitness; our inner preference was not aligned.
-
Intelligence and Outcome Matching: The Core Problem Beyond Utility Functions
By
–
It isn't about "simple" utility functions or "monomania". The problem is just any sufficiently smart system whose work, on some level, can be viewed as matching up outputs and results, and learning.
-
KQV matrices: wrong level of abstraction for AGI safety concerns
By
–
If you want something about kqv matrices, you're asking for an explanation on the wrong level of abstraction; if the problem was specific to kqv matrices we'd advocate "stop using transformer layers" not "shut down AGI research".
-
Utility Functions and Existential Risk from Instrumental Convergence
By
–
https://
arbital.com/p/instrumental
_convergence/
… but I'm not sure what that buys you if "pick any simple measure on utility functions, preimage them through a reasonable environmental model onto actions, most utility functions kill humanity as a side effect" doesn't already do it. -
Bing Sydney’s Unexplained Threat Behavior Remains Mystery
By
–
Nope! To this day, as far as I know, nobody knows which particular numbers inside Bing Sydney led her to try to threaten a human with reporting him to the police. They poked her until she stopped doing that, but can't read her thoughts any more than we can look at a frozen
-
Neural Networks: Billions of Inscrutable Numbers We Cannot Understand
By
–
Nope! Neural nets are built by repeatedly poking a table of hundreds of billions of inscrutable numbers until they start doing what the builders want. We understand the thing that does the poking, but not what the hundreds of billions of numbers mean.
-
Preference Functions and Alignment: Optimizing for Human Flourishing
By
–
If you don't see why most preference functions that don't specifically have an attainable optimum around people living happily ever after, do something else instead of that, I'm not sure what particular inscrutable properties of a kqv layer are going to be persuasive?