This approach fails for appointing benevolent human dictators to run our governments for us, because humans are smart enough to be pretend to be nicer than they are. So checking the apparent subservience of AIs isn't a reliable indicator once they're smart enough to fake that.
ETHICS
-
Bureaucracies Cannot Distinguish Quality Alignment Research Papers
By
–
I don't think that does it, because the bureaucracies doing the funding wouldn't know how to distinguish good alignment papers from bad alignment papers. It's possible that some progress could be made on AI interpretability this way; but I don't think that's enough.
-
Balancing Short-term and Long-term AI Risk Assessment
By
–
I focused on both short term and long term risks, and made sure both were represented
-
Optimization creates minds with desires through evolution
By
–
If you optimize hard enough over any open-ended problem, you get minds; minds that want things. Minds and wanting are effective ways of computing complicated answers. That's how humans, and human brains, came into existence just from evolution hill-climbing "how to reproduce".
-
General intelligence versus specialized AI systems in chess
By
–
You could imagine Magnus Carlsen pouring water onto the Stockfish computer, because he's much more general than Stockfish even if he's not as good as chess. Harder for a 10-year-old to do to Magnus Carlsen, who also knows more about the world beyond the chessboard.
-
Advanced AI Independence from Biological Systems and Proteins
By
–
That It doesn't need current protein systems (life forms) to do anything, including think for It. It can make different life forms. It can make better proteins. It can make life not out of amino acids. So It doesn't need wheat, cows, or humans.
-
Superior Intelligence Eliminates Need for Biological Cooperation
By
–
If something is not just smarter than humanity, but smarter than the cumulative optimization power of evolution, there's no humans It needs to trade with or animals It needs to raise or plants It needs to cultivate; anything life does, It can design better alternatives for.
-
Chess AI Legibility: Hardcoded Utility Functions in Narrow Domains
By
–
– Chess AIs are far more legible; we can point to particular bytes inside a chess-playing program and say what they mean.
– Relatedly, chess is a narrower domain, and AIs can solve it without being able to learn new domains.
– Thus we can literally hardcode the utility function. -
Chess as Metaphor: Superintelligence Operating Beyond Human Rule Constraints
By
–
The chess game is a metaphor for the full real-world environment, reality itself; and Magnus Carlsen is a metaphor for a superintelligence, which plays by fewer rules than any human because It has less need of artificial rules to organize Its own thinking.
-
Relevance of Goal-Oriented Minds to Environmental Problem-Solving
By
–
How is any of that relevant to stumbling over minds-with-goals as a way to solve hard environmental problems?