“ExpRL: Exploratory RL for LLM Mid-Training” Sparse reward RL works only when the base model can already find useful reasoning paths, but on hard problems it often gets no signal. This paper uses reference solutions as reward scaffolds instead of imitation targets, letting an
ExpRL uses reference solutions as reward scaffolds for exploratory RL
By
–
