RLHF gets far more powerful as models get bigger here you see GPT-3.5 vs GPT-4 performance. you'll notice before fine-tuning, GPT-4 is a little bit better. but the really big gains come from GPT-4 after RLHF fine-tuning—that's the real unlock RLHF scales with model size!
RLHF Scaling: Massive Performance Gains With Larger Models
By
–
