what if we stopped betting everything on one agent rollout? "The Unreasonable Effectiveness of Scaling Agents for Computer Use" Generates multiple trajectories in parallel & selects the best using "behavior narratives" 69.9% on OSWorld, nearly matching human-level 72%
Scaling Parallel Agents for Computer Use Tasks
By
–
