Gradient Guidance at Test Time of Flow Policies in Reinforcement Learning. This paper uses expressive flow policies for RL without making policy training fragile. Thus, they do not train the policy with updates.
By
–

Gradient Guidance at Test Time of Flow Policies in Reinforcement Learning. This paper uses expressive flow policies for RL without making policy training fragile. Thus, they do not train the policy with updates.