AI Dynamics

Global AI News Aggregator

About

SFT vs DPO: Finding the Sweet Spot in Model Training

In my experience, SFT is quite destructive, which is why I like DPO better. There might be a sweet spot with SFT and low LRs though. I haven't experimented with it that much tbh.

→ View original post on X — @maximelabonne