AI Dynamics

Global AI News Aggregator

About

Q-Learning by Inversion for Flow Policies in Offline RL

Q-Learning by Inversion. Flow policies should be excellent for offline RL, but their iterative generation of actions makes RL training painful. This paper transforms each flow step into an RL step, then uses inversion to recover

→ View original post on X — @askalphaxiv