Abstract:To further understand the underlying mechanism of various reinforcement learning (RL) algorithms and also to better use the optimization theory to make further progress in RL, many researchers begin to revisit the linear-quadratic regulator (LQR) problem, whose setting is simple and yet captures the characteristics of RL. Inspired by this, this work is concerned with the model-free design of stochastic LQR controller for linear systems subject to Gaussian noises, from the perspective of both RL and primal-dual optimization. From the RL perspective, we first develop a new model-free off-policy policy iteration (MF-OPPI) algorithm, in which the sampled data is repeatedly used for updating the policy to alleviate the data-hungry problem to some extent. We then provide a rigorous analysis for algorithm convergence by showing that the involved iterations are equivalent to the iterations in the classical policy iteration (PI) algorithm. From the perspective of optimization, we first reformulate the stochastic LQR problem at hand as a constrained non-convex optimization problem, which is shown to have strong duality. Then, to solve this non-convex optimization problem, we propose a model-based primal-dual (MB-PD) algorithm based on the properties of the resulting Karush-Kuhn-Tucker (KKT) conditions. We also give a model-free implementation for the MB-PD algorithm by solving a transformed dual feasibility condition. More importantly, we show that the dual and primal update steps in the MB-PD algorithm can be interpreted as the policy evaluation and policy improvement steps in the PI algorithm, respectively. Finally, we provide one simulation example to show the performance of the proposed algorithms.

Successive over Relaxation for Model-Free LQR Control of Discrete-Time Markov Jump Systems

Reinforcement Learning-Based $\mathcal{h}_{\infty }$ Control of 2-D Markov Jump Roesser Systems with Optimal Disturbance Attenuation

Robust H∞ Output Feedback Control for Markovian Jump Systems with Actuator Saturation

Robust H-infinity control for uncertain discrete-time Markovian jump systems with actuator saturation

Stochastic LQ optimal control for Markov jumping systems with multiplicative noise using reinforcement learning

Value iteration for LQR control of unknown stochastic-parameter linear systems

Cooperative Output Regulation Quadratic Control for Discrete-Time Heterogeneous Multiagent Markov Jump Systems

Model-free optimal controller for discrete-time Markovian jump linear systems: A Q-learning approach

Model-Free Design of Stochastic LQR Controller from Reinforcement Learning and Primal-Dual Optimization Perspective

A Successive over Relaxation Implicit Iterative Algorithm for Solving Stochastic Linear Systems with Markov Jumps

Finite-time L2−l∞ Tracking Control for Markov Jump Repeated Scalar Nonlinear Systems with Partly Usable Model Information

H∞$$ {h}_{\infty } $$ Optimal Output Tracking Control for Markov Jump Systems: A Reinforcement Learning‐based Approach

Model-free design of stochastic LQR controller from a primal–dual optimization perspective

Output Feedback Stabilization for Discrete Singular Systems with Random Abrupt Changes

Asynchronous Static Output-Feedback Control of Markovian Jump Linear Systems

Asynchronous Observer-Based Control for Exponential Stabilization of Markov Jump Systems.

Reinforcement Learning-Based Direct Adaptive Optimal Control of JLQ Model

Inverse Value Iteration and Q -Learning: Algorithms, Stability, and Robustness

Optimal control for continuous-time Markov jump singularly perturbed systems : A hybrid reinforcement learning scheme

Linear Quadratic Regulation and Stabilization of Discrete-Time Systems With Delay and Multiplicative Noise

Indefinite Robust Linear Quadratic Optimal Regulator for Discrete-Time Uncertain Singular Markov Jump Systems