Mitigating Catastrophic Forgetting in Continual RL via Certified Alignment
Maksim Anisimov ⋅ Matthew Wicker ⋅ Francesco Belardinelli
Abstract
Continual adaptation enables reinforcement learning (RL) agents to learn new capabilities, sometimes at the cost of forgetting behaviour in previous tasks. We propose CertAlign, a certified alignment framework that constrains downstream policy updates to a parameter subspace in which any policy provably retains the source-task behaviour. The method uses a sound differentiable surrogate, enabling projected gradient adaptation with provable alignment guarantees. Experiments in Frozen Lake and Lunar Lander environments show that CertAlign is the only adaptation method that preserves source-task behaviour both provably and empirically, while retaining non-trivial downstream plasticity.
Chat is not available.
Successful Page Load