AI Dynamics

Global AI News Aggregator

About

Comparing RL, SFT, and Self-Training for Model Training

Hi Lingesh, I added a section comparing RL, SFT and Self-Training to answer your question. SFT does not have a selection mechanism (R) and also it requires a supervised dataset. RL only requires the prompts 'o' as it generates the 'a's.

→ View original post on X — @nandodf