6. Scaling up RL This paper investigates how prolonged RL can enhance reasoning abilities in small language models across diverse domains.
Scaling Reinforcement Learning for Enhanced Reasoning in Small Models
By
–
By
–
6. Scaling up RL This paper investigates how prolonged RL can enhance reasoning abilities in small language models across diverse domains.