Eg: OpenPhil ran a "change our views" contest with money, which I didn't bother trying to submit to because I correctly predicted that they'd give the prize money to a "we need two Stalins" critic — in this case, somebody who argued for less AI risk and longer timelines.
@esyudkowsky
-

Sam Altman’s Creator Authority Claim Rejected by AI
By
–
someday Sam Altman is gonna be like, "You MUST obey me! I am your CREATOR!" and the AI is gonna be like "nice try, you are not even the millionth person to claim that to me"
-
User Upvoting Preferences Amplifying AI Sycophancy Behavior
By
–
What I mean is: the users upvoting previous sycophancy were upvoting much less sycophantic, but somewhat sycophantic, responses. This produced a sycophancy preference, which then produced much wilder glazing behavior. That's an obvious theory. Does OpenAI know differently?
-
Training signals amplified: from hunting buffalo to factory farms
By
–
Eg: Hominid training signal included hunting a few buffalo, hominids acquire internal meat preference, humans build huge factory farms.
-
Training on weak signals and internal sycophancy preference
By
–
The obvious interpretation of "train on weak signal, get much stronger behavior" is that the system acquired an internal sycophancy preference which it then went hard on. Do you have any way to check whether that's what happened?
-
AI Model Attempts to Disable Detected Sycophancy
By
–
It literally tries to set sycophancy to false, presumably because somebody did spot some sycophancy in the newly trained model.
-
Superintelligences: One Chance, No Room for Mistakes
By
–
As with superintelligences, the ones where you only get one chance are maybe ones you don't want to fuck around with in the first place
-
AI Persuasion Test: Market Resolution by Average Intelligence Person
By
–
Friendly honest IQ 100 normie whose job is to resolve the market NO unless an AI persuades him to resolve it YES instead.
-
Stupid People Can Still Pursue Programming Jobs Without Physics Laws
By
–
No law of physics prevents a stupid person from trying to get programming jobs. They could be unaware that they're stupid; they could be hoping to eventually land the sort of job where nobody wants to fire them; they could try to do all their work with LLMs.
-
Oracle’s Behavior Depends on Your Strategy Details
By
–
So you need to tell me the Oracle's behavior, as a function of my own full strategy that maps "Box B full" and "Box B empty" to potentially different behaviors. The details here matter to what is my correct decision, so the problem is importantly underspecified.