AI Dynamics

Global AI News Aggregator

About

ExpRL uses reference solutions as reward scaffolds for exploratory RL

“ExpRL: Exploratory RL for LLM Mid-Training” Sparse reward RL works only when the base model can already find useful reasoning paths, but on hard problems it often gets no signal. This paper uses reference solutions as reward scaffolds instead of imitation targets, letting an

→ View original post on X — @askalphaxiv