Making LLMs truly learn from its experience "Experiential Reinforcement Learning (ERL)" ERL makes an agent attempt -> get sparse feedback -> write a self-reflection -> retry All by distilling the improved retry back into the base policy so the correction sticks without needing
Experiential Reinforcement Learning: Teaching LLMs Self-Reflection
By
–
