Abstract:Batch reinforcement learning (RL) defines the task of learning from a fixed batch of data lacking exhaustive exploration. Worst-case optimality algorithms, which calibrate a value-function model class from logged experience and perform some type of pessimistic evaluation under the learned model, have emerged as a promising paradigm for batch RL. However, contemporary works on this stream have commonly overlooked the hierarchical decision-making structure hidden in the optimization landscape. In this paper, we adopt a game-theoretical viewpoint and model the policy learning diagram as a two-player general-sum game with a leader-follower structure. We propose a novel stochastic gradient-based learning algorithm: StackelbergLearner, in which the leader player updates according to the total derivative of its objective instead of the usual individual gradient, and the follower player makes individual updates and ensures transition-consistent pessimistic reasoning. The derived learning dynamic naturally lends StackelbergLearner to a game-theoretic interpretation and provides a convergence guarantee to differentiable Stackelberg equilibria. From a theoretical standpoint, we provide instance-dependent regret bounds with general function approximation, which shows that our algorithm can learn a best-effort policy that is able to compete against any comparator policy that is covered by batch data. Notably, our theoretical regret guarantees only require realizability without any data coverage and strong function approximation conditions, e.g., Bellman closedness, which is in contrast to prior works lacking such guarantees. Through comprehensive experiments, we find that our algorithm consistently performs as well or better as compared to state-of-the-art methods in batch RL benchmark and real-world datasets.

BATCH POLICY LEARNING IN AVERAGE REWARD MARKOV DECISION PROCESSES

Robust Batch Policy Learning in Markov Decision Processes

A Batch, Off-Policy, Actor-Critic Algorithm for Optimizing the Average Reward

Optimal Sample Complexity for Average Reward Markov Decision Processes

Study on an Average Reward Reinforcement Learning Algorithm

Stackelberg Batch Policy Learning

Provably Efficient Infinite-Horizon Average-Reward Reinforcement Learning with Linear Function Approximation

Finding Optimal Memoryless Policies of POMDPs under the Expected Average Reward Criterion

Near-Optimal Regret Bounds for Multi-batch Reinforcement Learning

Performance Bounds for Policy-Based Average Reward Reinforcement Learning Algorithms

Sharper Model-free Reinforcement Learning for Average-reward Markov Decision Processes

Online Reinforcement Learning in Markov Decision Process Using Linear Programming

Provable Policy Gradient Methods for Average-Reward Markov Potential Games

Relative Q-Learning for Average-Reward Markov Decision Processes with Continuous States

Variance-Reduced Policy Gradient Approaches for Infinite Horizon Average Reward Markov Decision Processes

Learning Infinite-Horizon Average-Reward Linear Mixture MDPs of Bounded Span

Stochastic first-order methods for average-reward Markov decision processes

A Structure-aware Online Learning Algorithm for Markov Decision Processes

Near-Optimal Offline Reinforcement Learning via Double Variance Reduction

Average Reward Adjusted Discounted Reinforcement Learning: Near-Blackwell-Optimal Policies for Real-World Applications

Provably Efficient Reinforcement Learning for Infinite-Horizon Average-Reward Linear MDPs