we all learn reinforcement learning in intro ML. it’s widely accepted as foundational, a pillar of the field yet outside of RLHF, everyone working on language models ignored RL for years folk wisdom said it “doesn’t work” what other fundamental tools are we neglecting?
Reinforcement Learning: Neglected Tool in Language Models
By
–