Abstract:In this paper, a novel adaptive dynamic programming (ADP)-based optimal control method is developed for discrete-time systems subject to constraints and disturbances. Particularly, a safe policy iteration scheme is designed to handle state and input constraints, including both hard and soft constraints, by converting the original policy improvement strategy into a constrained optimization problem with a prescribed state cost function. After that, an actor-critic-disturbance framework is introduced to address the constrained optimal control problem. The robust safety against disturbances is treated as a two-player zero-sum game, where the actor and disturbance neural networks are used to approximate the optimal control input and the disturbance policy, respectively. The convergence property of the proposed algorithm is analyzed, and the multi-step version of the proposed ADP scheme is derived based on this property. Simulation results are demonstrated and discussed to validate the effectiveness and performance of the proposed method. Note to Practitioners—Addressing constraints in optimal control problems is essential for guaranteeing the safe operation of controlled systems. However, conventional ADP algorithms struggle to simultaneously manage state and control input constraints during the search for the optimal solution. In real-world applications, another critical and common issue is the presence of external disturbances, where disturbances that cause the control object to deviate from the safe region must be constrained while seeking an optimal control policy. Bearing these factors in mind, this study presents a novel ADP scheme for solving optimal control problems of discrete-time systems, taking into account state and control constraints as well as the impact of disturbances. Moreover, the convergence analysis of the proposed SADP scheme is provided, offering a powerful theoretical foundation for guaranteeing the safety and feasibility of the controlled system during operation.

Costate-Supplement ADP for Model-Free Optimal Control of Discrete-Time Nonlinear Systems

Model-free Adaptive Dynamic Programming for Optimal Control of Discrete-time Affine Nonlinear System

ADP-Based Optimal Control for Discrete-Time Systems With Safe Constraints and Disturbances

Twin Deterministic Policy Gradient Adaptive Dynamic Programming for Optimal Control of Affine Nonlinear Discrete-time Systems

Data-Driven Event-Triggered Adaptive Dynamic Programming Control for Nonlinear Systems with Input Saturation.

Adaptive Dynamic Programming for Nonaffine Nonlinear Optimal Control Problem with State Constraints

Adaptive dynamic programming for optimal control of discrete‐time nonlinear system with state constraints based on control barrier function

Model-Free Incremental Adaptive Dynamic Programming Based Approximate Robust Optimal Regulation

Discrete‐Time Optimal Control of State‐Constrained Nonlinear Systems Using Approximate Dynamic Programming

Functional Nonlinear Model Predictive Control Based on Adaptive Dynamic Programming

Robust ADP Design for Continuous-Time Nonlinear Systems with Output Constraints

Output Constrained Adaptive Dynamic Programming For Continuous-Time Nonlinear Systems

Online optimal consensus control of unknown linear multi-agent systems via time-based adaptive dynamic programming

Adaptive Dynamic Programming and Data-Driven Cooperative Optimal Output Regulation with Adaptive Observers

ADP-Based Decentralized Controller Design for Nonlinear Time-Delay Interconnected Systems

Online and model-free supplementary learning control based on approximate dynamic programming

Optimal control of nonlinear system based on deterministic policy gradient with eligibility traces

Adaptive Multi-Step Evaluation Design With Stability Guarantee for Discrete-Time Optimal Learning Control

Novel iterative neural dynamic programming for data-based approximate optimal control design

Approximate Dynamic Programming for Constrained Piecewise Affine Systems with Stability and Safety Guarantees

Robust ADP-based solution of a class of nonlinear multi-agent systems with input saturation and collision avoidance constraints