Hi Lingesh, I added a section comparing RL, SFT and Self-Training to answer your question. SFT does not have a selection mechanism (R) and also it requires a supervised dataset. RL only requires the prompts 'o' as it generates the 'a's.
Comparing RL, SFT, and Self-Training for Model Training
By
–