AI Dynamics

Global AI News Aggregator

About

Gradient guidance at test time of flow policies in RL

Gradient Guidance at Test Time of Flow Policies in Reinforcement Learning. This paper uses expressive flow policies for RL without making policy training fragile. Thus, they do not train the policy with updates.

→ View original post on X — @askalphaxiv