AI Dynamics

Global AI News Aggregator

About

VIMPO derives value function from policy to improve GRPO

"VIMPO: Value-Implicit Policy Optimization for LLMs" While GRPO is simple because it avoids a critic, it still gives every token in a reasoning trace the same reward signal. This paper tries to get the best of both worlds by deriving the value function from the policy itself

→ View original post on X — @askalphaxiv