Abstract:H{\infty} control of nonlinear continuous-time system depends on the solution of the Hamilton-Jacobi-Isaacs (HJI) equation, which has been proved impossible to obtain a closed-form solution due to the nonlinearity of HJI equation. In order to solve HJI equation, many iterative algorithms were proposed, and most of the algorithms were essentially Newton method when the fixed-point equation was constructed in a Banach space. Newton method is a local optimization method, it has small convergence region and needs the initial guess to be sufficiently close to the solution. Whereas damped Newton method enhances the robustness with respect to initial condition and has larger convergence region. In this paper, a novel reinforcement learning method which is named {\alpha}-policy iteration ({\alpha}-PI) is introduced for solving HJI equation. First, by constructing a damped Newton iteration operator equation, a generalized Bellman equation (GBE) is obtained. The GBE is an extension of bellman equation. And then, by iterating on the GBE, an on-policy {\alpha}-PI reinforcement learning method without using knowledge regarding to the system internal dynamics is proposed. Third, based on the on-policy {\alpha}-PI reinforcement learning method, we develop an off-policy {\alpha}-PI reinforcement learning method without requiring any knowledge of the system dynamics. Finally, the neural-network based adaptive critic implementation schemes of on-policy and off-policy {\alpha}-PI algorithms are derived respectively, and the batch least-squares method is used for calculating the weight parameters of neural networks. The effectiveness of the off-policy {\alpha}-PI algorithm is verified through computer simulation.

Online Adaptive Policy Learning Algorithm for H-Infinity State Feedback Control of Unknown Affine Nonlinear Discrete-Time Systems

Adaptive Optimal Control for a Class of Continuous-Time Affine Nonlinear Systems with Unknown Internal Dynamics

Reinforcement Learning Controller Design for Affine Nonlinear Discrete-Time Systems Using Online Approximators

Model-free Adaptive Dynamic Programming for Optimal Control of Discrete-time Affine Nonlinear System

Online reinforcement learning control of unknown nonaffine nonlinear discrete time systems

H<inf>&#x221E;</inf> Control of Unknown Discrete-Time Nonlinear Systems with Control Constraints Using Adaptive Dynamic Programming

Online Approximate Solution of HJI Equation for Unknown Constrained-Input Nonlinear Continuous-Time Systems.

Online Reinforcement Learning-based Neural Network Controller Design for Affine Nonlinear Discrete-time Systems.

Heuristic dynamic programming-based H∞ control of discrete-time affine nonlinear systems

Online Adaptive Optimal Control for Continuous-Time Nonlinear Systems with Completely Unknown Dynamics.

A Novel Policy Iteration Algorithm for Nonlinear Continuous-Time H$\infty$ Control Problem

Discrete-Time Adaptive Iterative Learning Control for High-Order Nonlinear Systems with Unknown Control Directions

Online Off-Policy Reinforcement Learning for Optimal Control of Unknown Nonlinear Systems Using Neural Networks

Neural network-based finite-horizon optimal control of uncertain affine nonlinear discrete-time systems

A Parallel Framework of Adaptive Dynamic Programming Algorithm with Off-Policy Learning.

Off-Policy Reinforcement Learning for $ H_\infty $ Control Design

Twin Deterministic Policy Gradient Adaptive Dynamic Programming for Optimal Control of Affine Nonlinear Discrete-time Systems

Policy-Iteration-Based Finite-Horizon Approximate Dynamic Programming for Continuous-Time Nonlinear Optimal Control

Data-driven iterative adaptive dynamic programming algorithm for approximate optimal control of unknown nonlinear systems

Indirect Adaptive Fuzzy-Regulated Optimal Control for Unknown Continuous-Time Nonlinear Systems.

Adaptive dynamic programming-based algorithm for infinite-horizon linear quadratic stochastic optimal control problems