I suspect that LLMs are mostly a huge mush of policies, rather than an engine with updating beliefs and stable preferences. If so we cannot sensibly ask of Claude whether it was fooled into doing something that it didn't want, or if we found an outcome it didn't dislike.
@esyudkowsky
-

LLM Alignment Difficulty Often Underestimated by Non-Experts
By
–
Your occasional reminder that people who don't understand the state of play in LLMs sometimes brag about how easy they are to align. (Keeping in mind that current "AI safety" is corporate brand safety; that is what they are trying to do, so that failure is what is informative.)
-

Agreement on Not Going Extinct in Dumbest Ways Matters
By
–
Actually, this symbolic gesture strikes me as extremely important. It's a big deal to have agreement in principle on not going extinct in the dumbest ways — even if they haven't identified all the worst dangers, yet. My gratitude to anyone who worked on either side of this.
-

China Understands Mutual Interest in AI Existential Risk
By
–
China is perfectly capable of seeing our common interest in not going extinct. The claim otherwise is truth-uncaring bullshit by AI executives trying to avoid regulation.
-
OpenAI’s Value Assessment and Clarification
By
–
I mean, not *most* valuable, but yes. (This is about OpenAI.)
-

LLMs Excel at Writing RPG Flavor Text with Mood and Style
By
–
Current LLMs are surprisingly good if you ask them to write RPG flavor text — it doesn't have to be logical or creative, it just needs mood and style; and *that*, LLMs can nail.
-
LLM Whisperer: Psychotic Tendencies and Pattern Recognition Skills
By
–
So there's a made-up speculation I have, which goes like this: maybe LLM Whisperer is a job where psychotic leanings are rewarded by skill acquisition. Let's say you read some text, and you form a theory: This text is trying to communicate a secret message to you, and this x.com/ESYudkowsky/st…
-
Concern about psychotic hypotheses about LLMs driving people insane
By
–
Entirely separately: I have a relatively unimportant concern about whether there's going to be a class of people who make up psychotic hypotheses about LLMs, continually find apparent evidence for those hypotheses as they keep talking to the LLM, and are driven more insane by
-
LLM Whisperer: Psychotic Tendencies and Pattern Recognition Skills
By
–
So there's a made-up speculation I have, which goes like this: maybe LLM Whisperer is a job where psychotic leanings are rewarded by skill acquisition. Let's say you read some text, and you form a theory: This text is trying to communicate a secret message to you, and this
-
Why CEV Rather Than Something Else? Merits of the Proposal
By
–
Why CEV rather than something else? What is its merit as a proposal?