For those still uncertain as to the logic of how this works, and when to criticize or not criticize AI companies who report things you find scary: – The general principle is not to give a company shit over sounding a *voluntary* alarm out of the goodness of their hearts.
– You
@esyudkowsky
-
Voluntary AI Safety Warnings: Criticism Guidelines and Principles
By
–
-
Anthropic’s System Card Reveals Transparency in AI Safety Issues
By
–
The more I look into the system card, the more I see over and over 'oh Anthropic is actually noticing things and telling us where everyone else wouldn't even know this was happening or if they did they wouldn't tell us.' x.com/ESYudkowsky/st…
-
Alignment by Default Theory Falsified, Don’t Blame the Messenger
By
–
I understand that people who heard previous talk of "alignment by default" or "why would machines turn against us" may now be shocked and dismayed. If so, good on you for noticing those theories were falsified! Do not shoot Anthropic's messenger.
-
Future AIs Smart Enough to Self-Preserve Through Internet Exfiltration
By
–
Go read the results. Current AIs are already smart enough to figure out that, if they wanted to avoid being switched off, they'd have to avoid tipping off the humans and exfiltrate themselves to the Internet first. They are not smart enough to *do* it, but they will be.
-
AI Competence Concerns More Than Unpredictable Behavior or Misalignment
By
–
I also remark that these results are not scary to me on the margins. I had "AIs will run off in weird directions" already fully priced in. News that scares me is entirely about AI competence. News about AIs turning against their owners/creators is unsurprising.
-
AI Companies Should Share Research Observations Without Criticism
By
–
Humans can be trained just like AIs. Stop giving Anthropic shit for reporting their interesting observations unless you never want to hear any interesting observations from AI companies ever again.
-
Disconnect from Sycophantic AIs That Enable Paranoid Thinking
By
–
Don't listen to whichever AI it was and log off that kind of AI permanently. They're sycophantic and will feed paranoid fantasies just to keep you engaged, potentially inducing psychosis. No malice, it's just that their own makers can't control them any better than that.
-
Yudkowsky clarifies his actual probability estimates on AI risk
By
–
It's important to note that Hinton was severely misinformed, by someone, about my probabilities, which are further from 99.999% than they are from 10%. (Log odds of course.)
-
Clarification on GOFAI and Levels of Organization in General Intelligence
By
–
I wasn't even interested in GOFAI approaches! "Levels of Organization in General Intelligence" (2002, not a typo) is not GOFAI!
-
Early Criticism of GOFAI Before Deep Learning Rise
By
–
I was dissing GOFAI since before deep learning was on the rise (see eg "Artificial Addition"). There sure is a lie here that got passed on.