This is so neat! Dynamic Fine-Tuning (DFT) reweights the SFT loss by the model's own token probability, which creates a feedback loop. So they added forward KL to penalize any token the base finds likely, but the policy has pushed toward zero probability. DFT and forward KL
Dynamic Fine-Tuning (DFT) combines SFT loss with forward KL penalty
By
–
