7/ Fine-Grained RLHF – trains LMs with fine-grained human feedback; instead of using overall preference, more explicit feedback is provided at the segment level which helps to improve efficacy on long-form question answering and reduces toxicity.
Fine-Grained RLHF Improves LM Training and Safety
By
–
