Abstract:To further understand the underlying mechanism of various reinforcement learning (RL) algorithms and also to better use the optimization theory to make further progress in RL, many researchers begin to revisit the linear-quadratic regulator (LQR) problem, whose setting is simple and yet captures the characteristics of RL. Inspired by this, this work is concerned with the model-free design of stochastic LQR controller for linear systems subject to Gaussian noises, from the perspective of both RL and primal-dual optimization. From the RL perspective, we first develop a new model-free off-policy policy iteration (MF-OPPI) algorithm, in which the sampled data is repeatedly used for updating the policy to alleviate the data-hungry problem to some extent. We then provide a rigorous analysis for algorithm convergence by showing that the involved iterations are equivalent to the iterations in the classical policy iteration (PI) algorithm. From the perspective of optimization, we first reformulate the stochastic LQR problem at hand as a constrained non-convex optimization problem, which is shown to have strong duality. Then, to solve this non-convex optimization problem, we propose a model-based primal-dual (MB-PD) algorithm based on the properties of the resulting Karush-Kuhn-Tucker (KKT) conditions. We also give a model-free implementation for the MB-PD algorithm by solving a transformed dual feasibility condition. More importantly, we show that the dual and primal update steps in the MB-PD algorithm can be interpreted as the policy evaluation and policy improvement steps in the PI algorithm, respectively. Finally, we provide one simulation example to show the performance of the proposed algorithms.

Inverse Reinforcement Learning for Discrete-Time Linear Quadratic Systems

Average Cost Optimal Control of Stochastic Systems Using Reinforcement Learning

Reinforcement Learning for Inverse Linear-quadratic Dynamic Non-cooperative Games

Inverse reinforcement learning by expert imitation for the stochastic linear-quadratic optimal control problem

Reinforcement Learning for Inverse Non-Cooperative Linear-Quadratic Output-feedback Differential Games

Control of Nonaffine Nonlinear Discrete-Time Systems Using Reinforcement-Learning-Based Linearly Parameterized Neural Networks

Reinforcement Learning-Based Control for Nonlinear Discrete-Time Systems with Unknown Control Directions and Control Constraints

Online reinforcement learning control of unknown nonaffine nonlinear discrete time systems

Inverse Value Iteration and Q -Learning: Algorithms, Stability, and Robustness

Inverse Reinforcement Q-Learning Through Expert Imitation for Discrete-Time Systems

Quadratic Tracking Control of Linear Stochastic Systems with Unknown Dynamics Using Average Off-Policy Q-Learning Method

Inverse Reinforcement Learning for Identification of Linear-Quadratic Zero-Sum Differential Games

Model-Free Design of Stochastic LQR Controller from Reinforcement Learning and Primal-Dual Optimization Perspective

Inverse optimal control for averaged cost per stage linear quadratic regulators

Inverse reinforcement learning methods for linear differential games

Model-free design of stochastic LQR controller from a primal–dual optimization perspective

Reinforcement Learning-Based Direct Adaptive Optimal Control of JLQ Model

Reinforcement Learning for a Discrete-Time Linear-Quadratic Control Problem with an Application

Inverse linear-quadratic nonzero-sum differential games

Actively Learning Reinforcement Learning: A Stochastic Optimal Control Approach