Well-Posed KL-Regularized Control via Wasserstein and Kalman–Wasserstein KL Divergences
Abstract
Kullback-Leibler (KL) divergence regularization is widely used in reinforcement learning, but it becomes infinite under support mismatch and can degenerate in low-noise regimes. Using a unified information-geometric framework, we introduce KL analogs by replacing the Fisher–Rao geometry in the dynamical formulation of the KL with transport-based geometries, and derive closed-form expressions for common distribution families. Between elliptic distributions, these divergences remain finite for degenerating equal covariances and yield a geometric interpretation of regularization heuristics used in Kalman ensemble methods. We demonstrate the utility of these divergences in KL-regularized optimal control. In the fully tractable setting of linear time-invariant systems with Gaussian process noise, the classical KL reduces to a quadratic control penalty that becomes singular as process noise vanishes. Our variants remove this singularity and yield well-posed problems. On a double integrator and a cart-pole example, the resulting controls preserve nontrivial feedback and achieve better closed-loop performance.
Lay Summary
Many modern machine-learning and control methods use the Kullback–Leibler divergence to measure how far one probability distribution is from another. However, this standard measure can behave poorly when distributions do not overlap or when the system noise becomes very small. In such cases, KL-based regularization may become infinite or force the controller to act almost trivially. This paper develops new KL-like divergences that take the geometry of the underlying state space into account. Instead of measuring differences only statistically, these divergences also consider how far states are from each other in space. The resulting Wasserstein and Kalman–Wasserstein KL divergences remain finite in important low-noise regimes. The paper shows that these divergences admit explicit formulas for common distribution families and applies them to linear-quadratic control problems. In simple benchmark systems, the proposed regularizers avoid the degeneracy of classical KL regularization and lead to more meaningful feedback control.