Convergence Guarantees for Differentiable Optimization-based Control Policy

Yuexin Bian,Jie Feng,Yuanyuan Shi
2024-11-12
Abstract:Effective control of real-world systems necessitates the development of controllers that are not only performant but also interpretable. To this end, the field has seen a surge in model-based control policies, which first leverage historical data to learn system cost and dynamics, and then utilize the learned models for control. However, due to this decoupling, model-based control policies fall short when deployed in optimal control settings and lack convergence guarantees for achieving optimality. In this paper, we present DiffOP, a Differentiable Optimization-based Policy for optimal control. In the proposed framework, control actions are derived by solving an optimization, where the control cost and system's dynamics can be parameterized as neural networks. The key idea of DiffOP, inspired by differentiable optimization techniques, is to jointly learn the control policy using both policy gradients and optimization gradients, while utilizing actual cost feedback during system interaction. Further, this study presents the first theoretical analysis of the convergence rates and sample complexity for learning the optimization control policy with a policy gradient approach.
Systems and Control
What problem does this paper attempt to address?