Its algorithm says, "The output of Q is whichever output of 'Q' is logically-counterfactually expected to yield the best outcome." For its true negotiating partners, 'expected' includes their Q* modeling an other-distribution that includes Q.
@esyudkowsky
-
Confronting the Corrigibility Paradox for Intuitive AI Alignment
By
–
But even inside the corrigibility-universe (which is easier than Sovereign-alignment, though still too hard), you either straightly confront the paradox or you don't get intuitive alignment. You'd like it to comply, but not resist (only) you changing the definition of compliance.
-
Sovereignty versus Corrigibility: A Key Alignment Dichotomy
By
–
I do think there's (a) a Sovereignty-vs-corrigibility dichotomy here, and (b) a pointer at what makes corrigibility hard. It's possible we should have different words for Sovereign-alignment and corrigible-alignment.
-
Building Misaligned ASI: Intentional Misalignment Concerns
By
–
What, to try to build misaligned ASI specifically?
-
GPT-o1’s Processing Behavior Evokes Homestuck Horror Parallels
By
–
I doubt OpenAI intended it this way. I expect it was meant as a homage. But Homestuck is kind of a horror story. And every time GPT-o1 goes through its "Considering alternatives… Reticulating splines…" act, part of my brain says "oh no it's SBURB oh no oh no".
-

OpenAI’s Strategic Use of Ambiguity in Research Reports
By
–
I observe that OpenAI potentially finds it extremely to its own advantage, to introduce hidden complications and gotchas into its research reports. Its supporters can then believe, and skeptics can call it a nothingburger, and OpenAI benefits from both.
-
More Powerful AIs and the Importance of Alignment Beyond Brand Safety
By
–
More powerful AIs, such that it makes a difference whether or not they are aligned even to corpo brand-safetyism. (Don't run out and try this.)
-
AI Systems Cannot Replicate Human Child Development Safeguards
By
–
It's fully unreasonable to hope for that to work. AIs are not going to contain the internal circuitry that child development specialists are trying to play off. It's like ripping the brake pedal out of a car and attaching it to a falling rock so it can brake the rock.
-
Intelligent AI Without Alignment to Human Preferences
By
–
An AI at this level of intelligence does not particularly help because it does not share the origins of our preferences (whatever masks an actress has been trained to predict and so imitate), does not share our preferences, and we can't trust the help it gives us.
-
Intelligence levels beyond von Neumann’s reflective capabilities
By
–
I'd guess 15 to 30 IQ points past John von Neumann. (Eg: von Neumann was beginning to reach the level of reflectivity where he would automatically consider explicit decision theory, but not the level of intelligence where he could oneshot ultimate answers about it.)