In my experience, SFT is quite destructive, which is why I like DPO better. There might be a sweet spot with SFT and low LRs though. I haven't experimented with it that much tbh.
SFT vs DPO: Finding the Sweet Spot in Model Training
By
–
By
–
In my experience, SFT is quite destructive, which is why I like DPO better. There might be a sweet spot with SFT and low LRs though. I haven't experimented with it that much tbh.