AI Dynamics

Global AI News Aggregator

About

Direct Preference Optimization: Training LLMs Without RLHF

8/ Direct Preference Optimization – while helpful to train safe & useful LLMs, RLHF can be complex and often unstable; this work proposes an approach to finetune LMs by solving a classification problem on the human preferences data, with no RL required.

→ View original post on X — @dair_ai