Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use This paper introduces Step-Wise Reinforcement Learning (SWiRL), a method for improving multi-step reasoning and tool use in language models through synthetic data generation and offline RL. Problem: Standard RL
SWiRL: Synthetic Data and RL for Language Model Reasoning
By
–
