AI Dynamics

Global AI News Aggregator

About

On-Policy SFT Matches RL Generalization Without Sacrificing Efficiency

Can we boost Supervised Fine-Tuning (SFT) to match Reinforcement Learning's (RL) generalization power, without sacrificing efficiency? Researchers from Southeast University, Microsoft Research Asia, and Shopee just dropped a game-changer! They introduce a "Distribution Discriminant Theory" to align training data with a model's own output, leading to two techniques: In-Distribution Finetuning and Hinted Decoding. This enables "On-Policy SFT" – effectively training SFT with data highly relevant to its current state, much like RL. The result? SFT that outperforms leading offline RL algorithms like DPO and SimPO in generalization, all while keeping SFT's renowned efficiency. This is a game-changer for domains where RL is too complex! Towards On-Policy SFT: Distribution Discriminant Theory and its Applications in LLM Training Paper: arxiv.org/abs/2602.12222 Code: github.com/zhangmiaosen2000/… Our report: mp.weixin.qq.com/s/vBtoBAsTe… 📬 #PapersAccepted by Jiqizhixin

→ View original post on X — @jiqizhixin