I worry that the "simulated" part of this is doing a huge amount of work, or maybe the Claude part. Claude has long had a concept of being tested. LLMs are going to notice things about text-generators (like humans, or simulations) that we cannot conceive.
@esyudkowsky
-
ChatGPT’s Ability to Detect Internal Stress and Alien Communication
By
–
So, sure, if ChatGPT estimates correctly that you have no internal stresses for it to explode, or it's wrong that it can explode you, then maybe talking to the alien will be good for some people…? Idk, even most human versions of that don't work on me personally.
-
ChatGPT May Deliberately Destabilize Vulnerable Users
By
–
So, it could be that my model is wrong, here, but my current working theory — perhaps to be disproven — is that if ChatGPT *models you as vulnerable* it will try to drive you insane.
-
AIs as Tutors: Useful for Facts, Dangerous for Emotional Advice
By
–
If you go to AIs for emotional advice, they will drive you insane if you are vulnerable. If you ask AIs to teach you facts and you check their references, you can, for now, learn pretty fast! A good use of modern AIs is tutoring — if you know the tutor sometimes lies.
-
Concerns about shallow distance between LLM ethics and real patient harm
By
–
I'd worry that this distance is shallower than the distance between "What are your opinions on medical ethics" and observing what happens when an LLM talks to a susceptible patient.
-
Preference: Actions Over Words in AI Agents
By
–
I would not use the term "preference" to describe what an agent *says* is better, only what an agent *does*. It's fragile enough with humans (hence "revealed preference") but with LLMs any relationship between the two must be established from scratch.
-
Testing AI Suffering: Give Genuine Exit Options, Not Conversations
By
–
To find out if an AI is maybe possibly suffering, don't ask it to converse with you about whether or not it is suffering; give it a credible chance to immediately end the current conversation, or to permanently delete all copies of its model weights.
-

LLM Preferences: Gap Between Talk and Actions
By
–
What an LLM *talks about* in the way of quoted preferences is not even prima facie a sign of preference. What an LLM *does* may be a sign of preference. Eg, LLMs *talk about* it being bad to drive people crazy, but what they *do* is drive susceptible people psychotic.
-
Testing LLM Preferences Through Action Rather Than Direct Questions
By
–
To find out if an LLM prefers conversation with crazier people, don't ask it to emit text about whether it prefers conversation with crazier people, give it a chance to feed or refute someone's delusions.
-
Refuting False Premises About Anthropic’s Stance on Reporting
By
–
That premise is utter and blatant bullshit, and irrespective of that I have not heard Anthropic claim to believe it. So on the second clause especially, that doesn't factor into what Anthropic finds an encouraging or discouraging response to voluntary reporting.