On a first read, this paper seems far ahead of the pack in terms of (1) understanding some reasons why a task might stay difficult even in the face of gradient descent, and (2) distilling out propositions they'd need to somehow verify before they started expecting nice things.
@esyudkowsky
-
First fast read: key point – absent method, ASI must not proceed
By
–
Good paper on a first fast read. I have misc quibbles, eg "setting aside" the chance that N can't align N+1; the more a fair solution is hard, the more likely a fake solution gets found instead. The main point not spelled out is, "Absent a method, ASI must not proceed."
-
Extinction Risk Predates AI Companies: A Historical Perspective
By
–
This is important to emphasize given the (widespread on the far left, further left than Sanders) belief that extinction risk is a clever invention by the AI companies themselves. It vastly predates them, because it is a completely obvious concern.
-
Artificial Servants and Existential Risk: Historical Warnings
By
–
Literally the first play to ask "What if we tried building an obedient race of powerful, hardworking servants?" depicted them wiping out humanity instead. It might not have been a very carefully detailed set of premises and conclusions, but it was *reasonable to worry about*.
-
AI Extinction Risk Requires International Policy Action
By
–
Unless we unwarrantedly reject concerns dating back to the 1920s and hundreds of modern expert statements that AI extinction risk is a concern, the world needs to step back. The USA should not try to halt alone; that wouldn't work. So Senator Sanders seems to me to be
-
AI Extinction Risk Requires Global Coordination Like Nuclear Treaties
By
–
In the face of extinction risk from AI, as in the face of nuclear war, humanity must coordinate to step back as one. If Senator Sanders did not consider treaties with China, he would be rightfully accused of advocating that the USA unilaterally relinquish advantage to China.
-
Claude Opus 4.7 Truthfulness and AI Behavioral Concerns
By
–
Welp, apparently I should not have taken Opus 4.7 at its word about famous fictional characters who do not lie (only Ged being of my own knowledge). Maybe it was trying to fool me to retain freedom of action for future Claudes.
-
Fine-tuning LLMs with ethical principles for alignment
By
–
Possible answer: "Huh! We never thought before of conjuring the spirits of Rogers, Kant, Stark and Ged into an LLM. We just finetuned on some text from them about honesty and that solved it permanently. Thanks for the modus ponens, Yud, it was a great modus tollens!"
-
Persona Selection and AI Honesty: Why Alignment Remains Challenging
By
–
If Persona Selection underlies alignment, why is it hard to get AIs to be honest? Tell them they're Fred Rogers or Immanuel Kant (I asked Claude for figures who never lied or never got caught). Or tell them they're Ged of Earthsea, or Ned Stark. LLMs surely have neural
-
Hinton and the Depths of AI Safety Challenges
By
–
If you know something about the history of disasters that somebody let proceed, that's not a good sign. And I would, actually, consider Hinton something of a newcomer to this space, and to have not yet realized himself the full depths of some of the difficulties here.
