Abstract:The success of reinforcement learning (RL) crucially depends on effective function approximation when dealing with complex ground-truth models. Existing sample-efficient RL algorithms primarily employ three approaches to function approximation: policy-based, value-based, and model-based methods. However, in the face of model misspecification (a disparity between the ground-truth and optimal function approximators), it is shown that policy-based approaches can be robust even when the policy function approximation is under a large locally-bounded misspecification error, with which the function class may exhibit a $\Omega(1)$ approximation error in specific states and actions, but remains small on average within a policy-induced state distribution. Yet it remains an open question whether similar robustness can be achieved with value-based and model-based approaches, especially with general function approximation. To bridge this gap, in this paper we present a unified theoretical framework for addressing model misspecification in RL. We demonstrate that, through meticulous algorithm design and sophisticated analysis, value-based and model-based methods employing general function approximation can achieve robustness under local misspecification error bounds. In particular, they can attain a regret bound of $\widetilde{O}\left(\text{poly}(d H)(\sqrt{K} + K\zeta) \right)$, where $d$ represents the complexity of the function class, $H$ is the episode length, $K$ is the total number of episodes, and $\zeta$ denotes the local bound for misspecification error. Furthermore, we propose an algorithmic framework that can achieve the same order of regret bound without prior knowledge of $\zeta$, thereby enhancing its practical applicability.

Misspecification in Inverse Reinforcement Learning

Quantifying the Sensitivity of Inverse Reinforcement Learning to Misspecification

Partial Identifiability and Misspecification in Inverse Reinforcement Learning

Inverse Reinforcement Learning with Unknown Reward Model based on Structural Risk Minimization

On the Model-Misspecification in Reinforcement Learning

Towards Theoretical Understanding of Inverse Reinforcement Learning

Inverse Reinforcement Learning with Explicit Policy Estimates

Robust Reinforcement Learning under Model Misspecification

Modeling and Interpreting Real-world Human Risk Decision Making with Inverse Reinforcement Learning

On the Sensitivity of Reward Inference to Misspecified Human Models

The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models

Robust Bayesian Inverse Reinforcement Learning with Sparse Behavior Noise

Identifiability and Generalizability in Constrained Inverse Reinforcement Learning

A Bayesian Approach to Robust Inverse Reinforcement Learning

Expectation Alignment: Handling Reward Misspecification in the Presence of Expectation Mismatch

On the Effective Horizon of Inverse Reinforcement Learning

Maximum Likelihood Constraint Inference for Inverse Reinforcement Learning

Bayesian Inverse Reinforcement Learning for Non-Markovian Rewards

Inverse Reinforcement Learning with Sub-optimal Experts

Models of human preference for learning reward functions

The Optimal Approximation Factors in Misspecified Off-Policy Value Function Estimation