Thrilled to have the authors of Direct Preference Optimization answering questions on their paper! DPO offers a simple alternative to RLHF and has been hugely impactful, including being used to train Llama 3! Talk to the authors @archit_sharma97 and team directly to learn more!
Direct Preference Optimization Authors Discuss DPO Impact
By
–
