They are working on it; the problem is that they paid millions of dollars, notably OpenAI, to get RLHF databases where labelers tended to choose nice answers rather than useful ones. Whether it's human bias or a directive, we don't know, but it's been identified….
RLHF Bias: Millions for Nice but Useless Answers
By
–