AI Dynamics

Global AI News Aggregator

About

ReSeek: Self-Correcting Search Agents with Dense Rewards

Search agents could notice their own mistakes and fix their reasoning mid-trajectory. ReSeek enables exactly that. With a JUDGE action for on the fly self correction and dense rewards for factual correctness plus real utility, it trains agents that outperform SOTA on complex

→ View original post on X — @jiqizhixin