Great read: SFT generalizes worse than RL because the data is off-policy, not because of the SFT objective.
— Maxime Labonne (@maximelabonne) 2 octobre 2026
Rewriting the trajectories from an expert to look like what the base model writes can match or beat on-policy methods. Bonus: it also forgets less. https://t.co/90ovNBvRz8
Great read: SFT generalizes worse than RL because the data is off-policy, not because of the SFT objective. Rewriting the trajectories from an expert to look like what the base model writes can match or beat on-policy methods. Bonus: it also forgets less.