Iterative reflections for LLMs can outperform heavy RL? This paper shows that having the LLM reflects on its own trajectories, rewrite its own prompts, and evolve a diverse pool of candidates beats RL w/ GRPO so far on four reasoning tasks . 10% improv with 35x fewer rollouts!
Iterative LLM Reflections Outperform Heavy Reinforcement Learning
By
–
