Abstract:In this paper, a novel adaptive dynamic programming (ADP)-based optimal control method is developed for discrete-time systems subject to constraints and disturbances. Particularly, a safe policy iteration scheme is designed to handle state and input constraints, including both hard and soft constraints, by converting the original policy improvement strategy into a constrained optimization problem with a prescribed state cost function. After that, an actor-critic-disturbance framework is introduced to address the constrained optimal control problem. The robust safety against disturbances is treated as a two-player zero-sum game, where the actor and disturbance neural networks are used to approximate the optimal control input and the disturbance policy, respectively. The convergence property of the proposed algorithm is analyzed, and the multi-step version of the proposed ADP scheme is derived based on this property. Simulation results are demonstrated and discussed to validate the effectiveness and performance of the proposed method. Note to Practitioners—Addressing constraints in optimal control problems is essential for guaranteeing the safe operation of controlled systems. However, conventional ADP algorithms struggle to simultaneously manage state and control input constraints during the search for the optimal solution. In real-world applications, another critical and common issue is the presence of external disturbances, where disturbances that cause the control object to deviate from the safe region must be constrained while seeking an optimal control policy. Bearing these factors in mind, this study presents a novel ADP scheme for solving optimal control problems of discrete-time systems, taking into account state and control constraints as well as the impact of disturbances. Moreover, the convergence analysis of the proposed SADP scheme is provided, offering a powerful theoretical foundation for guaranteeing the safety and feasibility of the controlled system during operation.

Policy Optimization Adaptive Dynamic Programming for Optimal Control of Input-Affine Discrete-Time Nonlinear Systems.

Model-free Adaptive Dynamic Programming for Optimal Control of Discrete-time Affine Nonlinear System

Twin Deterministic Policy Gradient Adaptive Dynamic Programming for Optimal Control of Affine Nonlinear Discrete-time Systems

Bias-Policy Iteration-Based Adaptive Dynamic Programming for Optimal Control of Discrete-Time Nonlinear Systems

Adaptive Dynamic Programming for Optimal Control of Discrete-Time Nonlinear Systems with Trajectory-Based Initial Control Policy

Adaptive Dynamic Programming for Nonaffine Nonlinear Optimal Control Problem with State Constraints

A Parallel Framework of Adaptive Dynamic Programming Algorithm with Off-Policy Learning.

Optimal control of nonlinear system based on deterministic policy gradient with eligibility traces

Policy Gradient Adaptive Dynamic Programming for Model-Free Multi-Objective Optimal Control

ADP-Based Optimal Control for Discrete-Time Systems With Safe Constraints and Disturbances

Parallel Cross Entropy Policy Gradient Adaptive Dynamic Programming for Optimal Tracking Control of Discrete-Time Nonlinear Systems

Discrete‐Time Optimal Control of State‐Constrained Nonlinear Systems Using Approximate Dynamic Programming

Multi-step Heuristic Dynamic Programming for Optimal Control of Nonlinear Discrete-Time Systems.

Optimal Learning Output Tracking Control: A Model-Free Policy Optimization Method With Convergence Analysis

Policy-Iteration-Based Finite-Horizon Approximate Dynamic Programming for Continuous-Time Nonlinear Optimal Control

Modified λ-Policy Iteration Based Adaptive Dynamic Programming for Unknown Discrete-Time Linear Systems

Costate-Supplement ADP for Model-Free Optimal Control of Discrete-Time Nonlinear Systems

Modified $\lambda$-Policy Iteration Based Adaptive Dynamic Programming for Unknown Discrete-Time Linear Systems

Adaptive Optimal Control with Guaranteed Convergence Rate for Continuous-Time Linear Systems with Completely Unknown Dynamics.

Approximate Dynamic Programming for Constrained Piecewise Affine Systems with Stability and Safety Guarantees

Robust ADP Design for Continuous-Time Nonlinear Systems with Output Constraints