How far are we from agents that can self-generate world knowledge? The work proposes an outcome-based reward that measures how much an agent's self-generated world knowledge actually improves its task success rate. The external guidance is then removed at inference. Result: A
Agents That Self-Generate World Knowledge via Outcome-Based Rewards
By
–
