What if a robot could learn from its own successes, not just its failures? This research introduces SRPO, a method that lets vision-language-action models use their own best attempts as a reference to assign smarter rewards. This self-referential approach boosted a robot's
Robot Learning from Success: SRPO Method for Vision-Language Models
By
–
