AI Dynamics

Global AI News Aggregator

About

SFT generalizes worse than RL due to off-policy data

Great read: SFT generalizes worse than RL because the data is off-policy, not because of the SFT objective. Rewriting the trajectories from an expert to look like what the base model writes can match or beat on-policy methods. Bonus: it also forgets less.

→ View original post on X — @maximelabonne