Abstract:In reinforcement learning (RL) , one of the key components is policy evaluation, which aims to estimate the value function (i.e., expected long-term accumulated reward) of a policy. With a good policy evaluation method, the RL algorithms will estimate the value function more accurately and find a better policy. When the state space is large or continuous \emph{Gradient-based Temporal Difference(GTD)} policy evaluation algorithms with linear function approximation are widely used. Considering that the collection of the evaluation data is both time and reward consuming, a clear understanding of the finite sample performance of the policy evaluation algorithms is very important to reinforcement learning. Under the assumption that data are i.i.d. generated, previous work provided the finite sample analysis of the GTD algorithms with constant step size by converting them into convex-concave saddle point problems. However, it is well-known that, the data are generated from Markov processes rather than i.i.d. in RL problems.. In this paper, in the realistic Markov setting, we derive the finite sample bounds for the general convex-concave saddle point problems, and hence for the GTD algorithms. We have the following discussions based on our bounds. (1) With variants of step size, GTD algorithms converge. (2) The convergence rate is determined by the step size, with the mixing time of the Markov process as the coefficient. The faster the Markov processes mix, the faster the convergence. (3) We explain that the experience replay trick is effective by improving the mixing property of the Markov process. To the best of our knowledge, our analysis is the first to provide finite sample bounds for the GTD algorithms in Markov setting.

On Convergence Rate of MRetrace

Design from Policies: Conservative Test-Time Adaptation for Offline Policy Optimization

Convergence Rate of Primal-Dual Approach to Constrained Reinforcement Learning with Softmax Policy

Modified Retrace for Off-Policy Temporal Difference Learning.

An Off-Policy Trust Region Policy Optimization Method with Monotonic Improvement Guarantee for Deep Reinforcement Learning

Convergence Analysis of an Incremental Approach to Online Inverse Reinforcement Learning

A Convergent Off-Policy Temporal Difference Algorithm

Finite Sample Analysis of the GTD Policy Evaluation Algorithms in Markov Setting

Convergence Analysis of Asynchronous Stochastic Recursive Gradient Algorithms

On convergence rates of game theoretic reinforcement learning algorithms

A Variance Minimization Approach to Temporal-Difference Learning

On the Convergence of Reinforcement Learning with Monte Carlo Exploring Starts

On The Convergence Of Policy Iteration-Based Reinforcement Learning With Monte Carlo Policy Evaluation

Deadly triad matters for offline reinforcement learning

Global Convergence of Policy Gradient Methods in Reinforcement Learning, Games and Control

On the Convergence of Discounted Policy Gradient Methods

Diverse Policies Converge in Reward-free Markov Decision Processe

Trajectory-Oriented Policy Optimization with Sparse Rewards

Reusing Historical Trajectories in Natural Policy Gradient via Importance Sampling: Convergence and Convergence Rate

Non-asymptotic Convergence of Adam-type Reinforcement Learning Algorithms under Markovian Sampling

Fast Rates for the Regret of Offline Reinforcement Learning