7. AdaptThink This paper introduces AdaptThink, an RL framework designed to help reasoning models decide when to use detailed chain-of-thought reasoning (“Thinking”) versus directly producing an answer (“NoThinking”), based on task difficulty.
AdaptThink: RL Framework for Dynamic Reasoning Strategy Selection
By
–