7/ rStar – introduces self-play mutual reasoning to improve the reasoning capabilities of small language models without fine-tuning or superior models; MCTS is augmented with human-like reasoning actions, obtained from SLMs, to build richer reasoning trajectories…
rStar Enhances Small Language Models Reasoning Without Fine-tuning
By
–