Mixed‐Integer Optimal Control via Reinforcement Learning: A Case Study on Hybrid Electric Vehicle Energy Management

Jinming Xu,Nasser Lashgarian Azad,Yuan Lin
DOI: https://doi.org/10.1002/oca.3216
2024-10-17
Optimal Control Applications and Methods
Abstract:The proposed hybrid‐action reinforcement learning algorithm, TD3AQ, demonstrates promising potential in addressing mixed‐integer optimal control problems. Many optimal control problems require the simultaneous output of discrete and continuous control variables. These problems are typically formulated as mixed‐integer optimal control (MIOC) problems, which are challenging to solve due to the complexity of the solution space. Numerical methods such as branch‐and‐bound are computationally expensive and undesirable for real‐time control. This article proposes a novel hybrid‐action reinforcement learning (HARL) algorithm, twin delayed deep deterministic actor‐Q (TD3AQ), for MIOC problems. TD3AQ leverages actor‐critic and Q‐learning methods to manage discrete and continuous action spaces simultaneously. The proposed algorithm is evaluated on a plug‐in hybrid electric vehicle (PHEV) energy management problem, where real‐time control of the discrete variables, clutch engagement/disengagement and gear shift, and continuous variable, engine torque, is essential to maximize fuel economy while satisfying driving constraints. Simulation results show that TD3AQ achieves near‐optimal control, with only a 4.69% difference from dynamic programming (DP), and outperforms baseline reinforcement learning algorithms for hybrid action spaces. The sub‐millisecond execution time indicates potential applicability in other time‐critical scenarios, such as autonomous driving or robotic control.
automation & control systems,operations research & management science,mathematics, applied
What problem does this paper attempt to address?