Abstract:Batch reinforcement learning (RL) defines the task of learning from a fixed batch of data lacking exhaustive exploration. Worst-case optimality algorithms, which calibrate a value-function model class from logged experience and perform some type of pessimistic evaluation under the learned model, have emerged as a promising paradigm for batch RL. However, contemporary works on this stream have commonly overlooked the hierarchical decision-making structure hidden in the optimization landscape. In this paper, we adopt a game-theoretical viewpoint and model the policy learning diagram as a two-player general-sum game with a leader-follower structure. We propose a novel stochastic gradient-based learning algorithm: StackelbergLearner, in which the leader player updates according to the total derivative of its objective instead of the usual individual gradient, and the follower player makes individual updates and ensures transition-consistent pessimistic reasoning. The derived learning dynamic naturally lends StackelbergLearner to a game-theoretic interpretation and provides a convergence guarantee to differentiable Stackelberg equilibria. From a theoretical standpoint, we provide instance-dependent regret bounds with general function approximation, which shows that our algorithm can learn a best-effort policy that is able to compete against any comparator policy that is covered by batch data. Notably, our theoretical regret guarantees only require realizability without any data coverage and strong function approximation conditions, e.g., Bellman closedness, which is in contrast to prior works lacking such guarantees. Through comprehensive experiments, we find that our algorithm consistently performs as well or better as compared to state-of-the-art methods in batch RL benchmark and real-world datasets.

Q-learning with biased policy rules

Adaptive Order Q-learning

Adapting Double Q-Learning for Continuous Reinforcement Learning

Extending Q-learning to continuous and mixed strategy games based on spatial reciprocity

Multiagent Soft Q-Learning

Ensemble Bootstrapping for Q-Learning

Maxmin Q-learning: Controlling the Estimation Bias of Q-learning

Balanced Q-learning: Combining the Influence of Optimistic and Pessimistic Targets

Stackelberg Batch Policy Learning

Mildly Conservative Q-Learning for Offline Reinforcement Learning

Interaction state Q-learning promotes cooperation in the spatial prisoner's dilemma game

Mitigating Off-Policy Bias in Actor-Critic Methods with One-Step Q-learning: A Novel Correction Approach

On the Estimation Bias in Double Q-Learning

Status-quo policy gradient in Multi-Agent Reinforcement Learning

An Information-Theoretic Optimality Principle for Deep Reinforcement Learning

Biased Aggregation, Rollout, and Enhanced Policy Improvement for Reinforcement Learning

A Policy-Gradient Approach to Solving Imperfect-Information Games with Iterate Convergence

Generalized Individual Q-learning for Polymatrix Games with Partial Observations

Learning a Diffusion Model Policy from Rewards via Q-Score Matching

Deep Reinforcement Learning with Double Q-Learning

Best Possible Q-Learning