Large language model agents still struggle with multi-step reasoning, where one misstep can collapse the whole plan. While Process Reward Models (PRMs) aim to correct reasoning step-by-step using RL, they don’t scale well due to the high cost of evaluating tons of action
