Reinforcement Learning for Reasoning in Large Language Models with One Training Example This paper demonstrates that Reinforcement Learning with Verifiable Reward (RLVR) using just one training example (1-shot RLVR) can significantly enhance the mathematical reasoning abilities
Reinforcement Learning Enhances LLM Mathematical Reasoning
By
–
