Abstract:H{\infty} control of nonlinear continuous-time system depends on the solution of the Hamilton-Jacobi-Isaacs (HJI) equation, which has been proved impossible to obtain a closed-form solution due to the nonlinearity of HJI equation. In order to solve HJI equation, many iterative algorithms were proposed, and most of the algorithms were essentially Newton method when the fixed-point equation was constructed in a Banach space. Newton method is a local optimization method, it has small convergence region and needs the initial guess to be sufficiently close to the solution. Whereas damped Newton method enhances the robustness with respect to initial condition and has larger convergence region. In this paper, a novel reinforcement learning method which is named {\alpha}-policy iteration ({\alpha}-PI) is introduced for solving HJI equation. First, by constructing a damped Newton iteration operator equation, a generalized Bellman equation (GBE) is obtained. The GBE is an extension of bellman equation. And then, by iterating on the GBE, an on-policy {\alpha}-PI reinforcement learning method without using knowledge regarding to the system internal dynamics is proposed. Third, based on the on-policy {\alpha}-PI reinforcement learning method, we develop an off-policy {\alpha}-PI reinforcement learning method without requiring any knowledge of the system dynamics. Finally, the neural-network based adaptive critic implementation schemes of on-policy and off-policy {\alpha}-PI algorithms are derived respectively, and the batch least-squares method is used for calculating the weight parameters of neural networks. The effectiveness of the off-policy {\alpha}-PI algorithm is verified through computer simulation.

A New Continuous-Time Policy Iteration for Time-Varying Nonlinear Systems

Continuous-Time Policy Optimization.

Policy-Iteration-Based Finite-Horizon Approximate Dynamic Programming for Continuous-Time Nonlinear Optimal Control

Continuous-Time Distributed Policy Iteration for Multicontroller Nonlinear Systems

A Novel Policy Iteration Algorithm for Nonlinear Continuous-Time H$\infty$ Control Problem

Relaxed Actor-Critic with Convergence Guarantees for Continuous-Time Optimal Control of Nonlinear Systems.

Modified λ-Policy Iteration Based Adaptive Dynamic Programming for Unknown Discrete-Time Linear Systems

Ternary Policy Iteration Algorithm for Nonlinear Robust Control

Model-free Adaptive Dynamic Programming for Optimal Control of Discrete-time Affine Nonlinear System

Modified $\lambda$-Policy Iteration Based Adaptive Dynamic Programming for Unknown Discrete-Time Linear Systems

Two‐loop reinforcement learning algorithm for finite‐horizon optimal control of continuous‐time affine nonlinear systems

Bias-policy iteration based optimal control for unknown continuous-time linear periodic systems

Model-Free $δ$-Policy Iteration Based on Damped Newton Method for Nonlinear Continuous-Time H$\infty$ Tracking Control

AN OPEN-CLOSED-LOOP PI-TYPE ITERATIVE LEARNING CONTROL SCHEME FOR DISCRETE NONLINEAR TIME-VARYING SYSTEMS AND ITS CONVERGENCE

Adaptive Iterative Learning Control Mechanism for Nonlinear Systems Subject to High-Order Internal Model

Adaptive Iterative Learning Control for Discrete‐time Nonlinear Systems with Multiple Iteration‐varying High‐order Internal Models

Iteration-dependent High-order Internal Model Based Iterative Learning Control for Discrete-time Nonlinear Systems with Time-iteration-varying Parameter

Approximate Finite-Horizon Optimal Control with Policy Iteration

Adaptive Iterative Learning Control for Discrete-Time Systems with Non-Repetitiveness

Iteration-dependent High-order Internal Model Based Iterative Learning Control for Continuous-time Nonlinear Systems

On Iterative Learning Control for MIMO Nonlinear Systems in the Presence of Time-Iteration-varying Parameters