8/ Direct Preference Optimization – while helpful to train safe & useful LLMs, RLHF can be complex and often unstable; this work proposes an approach to finetune LMs by solving a classification problem on the human preferences data, with no RL required.https://t.co/DZ0GSarfuT
— DAIR.AI (@dair_ai) 4 juin 2023
8/ Direct Preference Optimization – while helpful to train safe & useful LLMs, RLHF can be complex and often unstable; this work proposes an approach to finetune LMs by solving a classification problem on the human preferences data, with no RL required.
