Open-ended reasoning is one of the hardest problems in reasoning LLMs rn. So in this paper, they aim to solve this by reverse-engineering plausible thought chains from good answers via a gradient-free search With DeepWriter-8B trained on this data outperforming top OS models!
Reverse-Engineering Thought Chains Improves LLM Reasoning Capabilities
By
–
