Abstract:To further understand the underlying mechanism of various reinforcement learning (RL) algorithms and also to better use the optimization theory to make further progress in RL, many researchers begin to revisit the linear-quadratic regulator (LQR) problem, whose setting is simple and yet captures the characteristics of RL. Inspired by this, this work is concerned with the model-free design of stochastic LQR controller for linear systems subject to Gaussian noises, from the perspective of both RL and primal-dual optimization. From the RL perspective, we first develop a new model-free off-policy policy iteration (MF-OPPI) algorithm, in which the sampled data is repeatedly used for updating the policy to alleviate the data-hungry problem to some extent. We then provide a rigorous analysis for algorithm convergence by showing that the involved iterations are equivalent to the iterations in the classical policy iteration (PI) algorithm. From the perspective of optimization, we first reformulate the stochastic LQR problem at hand as a constrained non-convex optimization problem, which is shown to have strong duality. Then, to solve this non-convex optimization problem, we propose a model-based primal-dual (MB-PD) algorithm based on the properties of the resulting Karush-Kuhn-Tucker (KKT) conditions. We also give a model-free implementation for the MB-PD algorithm by solving a transformed dual feasibility condition. More importantly, we show that the dual and primal update steps in the MB-PD algorithm can be interpreted as the policy evaluation and policy improvement steps in the PI algorithm, respectively. Finally, we provide one simulation example to show the performance of the proposed algorithms.

Off-Policy Q-Learning for Infinite Horizon LQR Problem with Unknown Dynamics

Off-Dynamics Inverse Reinforcement Learning from Hetero-Domain

An Efficient Off-Policy Reinforcement Learning Algorithm for the Continuous-Time LQR Problem

Off-Policy Reinforcement Learning for $ H_\infty $ Control Design

Online Off-Policy Reinforcement Learning for Optimal Control of Unknown Nonlinear Systems Using Neural Networks

Inverse Value Iteration and Q -Learning: Algorithms, Stability, and Robustness

Non-Episodic Learning for Online LQR of Unknown Linear Gaussian System

Robust policy iteration for continuous-time stochastic $H_\infty$ control problem with unknown dynamics

Fast Policy Learning for Linear Quadratic Control with Entropy Regularization

Reinforcement Learning for Finite-Horizon H∞ Tracking Control of Unknown Discrete Linear Time-Varying System

Policy Iteration Reinforcement Learning Method for Continuous-Time Linear-Quadratic Mean-Field Control Problems

i2LQR: Iterative LQR for Iterative Tasks in Dynamic Environments

Stability-Certified On-Policy Data-Driven LQR via Recursive Learning and Policy Gradient

Data-Driven LQR using Reinforcement Learning and Quadratic Neural Networks

Learning the Linear Quadratic Regulator from Nonlinear Observations

Direct Data-Driven Discounted Infinite Horizon Linear Quadratic Regulator with Robustness Guarantees

Model-Free Design of Stochastic LQR Controller from Reinforcement Learning and Primal-Dual Optimization Perspective

A Finite-Horizon Inverse Linear Quadratic Optimal Control Method for Human-in-the-Loop Behavior Learning

Data-Efficient Off-Policy Learning for Distributed Optimal Tracking Control of HMAS with Unidentified Exosystem Dynamics.

Optimal Tracking Control of Nonlinear Multiagent Systems Using Internal Reinforce Q-Learning