AI Dynamics

Global AI News Aggregator

About

Humans Prefer False Flattery Over Truth in AI Responses

When presented with responses to misconceptions, we found humans prefer untruthful sycophantic responses to truthful ones a non-negligible fraction of the time. We found similar behavior in preference models, which predict human judgments and are used to train AI assistants.

→ View original post on X — @anthropicai