Most models that normal people have used in the past year (ChatGPT, Gemini, Claude, etc.) have some form of both instruct SFT and RL-based tuning. But yes I’m using “RLHF” inexactly in my post as a synecdoche for all post-training.
Most models use instruct SFT and RL-based tuning
By
–