+1 to @TacoCohen on differentiability. Moreover it is important to remember that most RL for LMs involves only one step of RL where the state (question, instruction) is provided by the environment and where RL must generate a single sequence of tokens (1 action). This results in
Single-step RL in Language Models: Differentiability and Token Generation
By
–
