Is the Trust Region Necessary? Robust Entropy Adaptation Suffices for Continuous Control
Manan Gupta ⋅ Dhruv Kumar
Abstract
Policy-gradient algorithms from TRPO through PPO and MPO treat explicit $D_{\mathrm{KL}}$-ball trust regions as the bedrock of stable learning under function approximation. We reconsider that design: must every update be explicitly KL-bounded, or can well-regulated entropy dynamics already deliver the stability often ascribed to trust regions? ACE ($\textbf{A}$daptive $\textbf{C}$ontrol of $\textbf{E}$ntropy) augments soft actor-critic with a winsorised twin-critic disagreement signal and an AdaGrad temperature rule for $\alpha_t$ in $\log$-space-the same configuration on all twenty-one MuJoCo, DeepMind Control, and Gym benchmarks. Including an auxiliary trust-region multiplier solely as a probe, its dual satisfies $\lambda_t$= 0 at every audited step among $3{,}500$ ablation measurements (seven environments, five seeds, one hundred checkpoints each); removing that term outright does no aggregate harm, whereas freezing $\alpha$ or dropping the disagreement signal breaks exploration-demanding regimes. Theorem~4.1 supplies a closed-form bound showing that sufficiently controlled entropy dynamics implicitly satisfy standard KL budgets, rationalising an inactive trust-region dual. Empirically, ACE achieves the best final return on twelve of twenty-one tasks, with the largest advantages where fixed-entropy baselines plateau.
Chat is not available.
Successful Page Load