Disentangled Differentiable Model Predictive Control for Data-efficient and Interpretable Imitation Learning
Abstract
Efficient imitation of expert behaviors in high-dimensional continuous control remains a fundamental challenge, particularly when balancing physical safety with adaptation to varying task environments. In this paper, we propose a framework that treats expert behavior as a dynamic composition of learnable control primitives---interpretable cost function components within a differentiable Model Predictive Control (MPC) layer. By reformulating imitation learning from a pure black-box regression into a structured grey-box optimization, our model disentangles complex expert strategies into shared strategic bases and a context-aware gating network. This gating mechanism, conditioned via Feature-wise Linear Modulation (FiLM), integrates temporal motion history with environmental context to dynamically modulate control primitive activations, ensuring seamless strategic transitions across varying geometries. We validate our approach on high-speed autonomous racing benchmarks, where the framework demonstrates superior fidelity and transparency. Notably, the architecture enables rapid adaptation to out-of-distribution contexts, achieving near-expert performance on unseen geometries via few-shot refinement with only a single lap of data. These results highlight the efficiency of structured differentiable layers in distilling robust, interpretable decision-making policies from offline datasets.