Abstract:In reinforcement learning (RL) , one of the key components is policy evaluation, which aims to estimate the value function (i.e., expected long-term accumulated reward) of a policy. With a good policy evaluation method, the RL algorithms will estimate the value function more accurately and find a better policy. When the state space is large or continuous \emph{Gradient-based Temporal Difference(GTD)} policy evaluation algorithms with linear function approximation are widely used. Considering that the collection of the evaluation data is both time and reward consuming, a clear understanding of the finite sample performance of the policy evaluation algorithms is very important to reinforcement learning. Under the assumption that data are i.i.d. generated, previous work provided the finite sample analysis of the GTD algorithms with constant step size by converting them into convex-concave saddle point problems. However, it is well-known that, the data are generated from Markov processes rather than i.i.d. in RL problems.. In this paper, in the realistic Markov setting, we derive the finite sample bounds for the general convex-concave saddle point problems, and hence for the GTD algorithms. We have the following discussions based on our bounds. (1) With variants of step size, GTD algorithms converge. (2) The convergence rate is determined by the step size, with the mixing time of the Markov process as the coefficient. The faster the Markov processes mix, the faster the convergence. (3) We explain that the experience replay trick is effective by improving the mixing property of the Markov process. To the best of our knowledge, our analysis is the first to provide finite sample bounds for the GTD algorithms in Markov setting.

Recursive Approaches for Single Sample Path Based Markov Reward Processes

Single Sample Path-Based Optimization of Markov Chains

A State Aggregation Approach to Singularly Perturbed Markov Reward Processes

Continuous-Time Markov Decision Process With Average Reward: Using Reinforcement Learning Method

A Basic Formula for Performance Gradient Estimation of Semi-Markov Decision Processes

Two Time-Scale Gradient Approximation Algorithm For Adaptive Markov Reward Processes

Simulation-Based optimization of singularly perturbed markov reward processes with states aggregation

On-Line Policy Gradient Estimation with Multi-Step Sampling.

Infinite-horizon gradient estimation for semi-Markov decision processes

Optimal Control of Ergodic Continuous-Time Markov Chains with Average Sample-Path Rewards

An improvement of policy gradient estimation algorithms

Loop Estimator for Discounted Values in Markov Reward Processes

Finite Sample Analysis of the GTD Policy Evaluation Algorithms in Markov Setting

Simulation Optimization Algorithm for SMDPs with Parameterized Randomized Stationary Policies

Risk-sensitive Markov Decision Process and Learning under General Utility Functions

Markov Decision Processes with State-Dependent Discount Factors and Unbounded Rewards/costs.

Thompson Sampling for Infinite-Horizon Discounted Decision Processes

Online Reinforcement Learning via Posterior Sampling of Policy

Improving Sample Efficiency of Model-Free Algorithms for Zero-Sum Markov Games

Continuous Time Markov Decision Processes with Expected Discounted Total Rewards

Convergence Rates of Accelerated Markov Gradient Descent with Applications in Reinforcement Learning