Reparameterization Flow Policy Optimization
Abstract
Lay Summary
Reparameterization Policy Gradient (RPG) achieves high sample efficiency by backpropagating gradients through differentiable dynamics, but it mostly relies on Gaussian policies. We observe that flow policies, which generate actions via differentiable ODE integration, are a natural structural match for RPG. We propose Reparameterization Flow Policy Optimization (RFO), which backpropagates jointly through the flow generation process and the environment dynamics, unlocking the sample efficiency of RPG for flow policies without ever computing intractable log-likelihoods. Two lightweight regularizers handle stability and exploration. Across diverse rigid- and soft-body locomotion and manipulation tasks, RFO consistently beats strong baselines, nearly doubling the reward of the state-of-the-art baseline on a challenging soft-body quadruped task.