Thanks for these thoughts. Very helpful. RL does indeed enable us to generalise better. I fully agree. In principle, RL results in causal knowledge whereas SFT only on association. But I feel the question is still relevant because we still rely heavily on (self-)supervised
Reinforcement Learning Enables Causal Knowledge Beyond Supervised Fine-Tuning
By
–