Our latest policy brief highlights a quantitative framework that evaluates bias in language models. One key takeaway? Language models fine-tuned with human feedback were less representative of public opinion than models that were not fine-tuned. Read more: https://
stanford.io/3LXLCWI
Language Models Fine-Tuned Feedback Show Less Public Opinion Bias
By
–
