"VIMPO: Value-Implicit Policy Optimization for LLMs" While GRPO is simple because it avoids a critic, it still gives every token in a reasoning trace the same reward signal. This paper tries to get the best of both worlds by deriving the value function from the policy itself
