Why can't smaller language models match larger ones on reasoning? Researchers from CUHK Shenzhen, Alibaba Group, and Westlake University introduce G²RPO-A. It adaptively feeds correct reasoning steps into training, dynamically adjusting guidance as the model improves. On math
