Abstract:In this work, we consider the problem of model selection for deep reinforcement learning (RL) in real-world environments. Typically, the performance of deep RL algorithms is evaluated via on-policy interactions with the target environment. However, comparing models in a real-world environment for the purposes of early stopping or hyperparameter tuning is costly and often practically infeasible. This leads us to examine off-policy policy evaluation (OPE) in such settings. We focus on OPE for value-based methods, which are of particular interest in deep RL, with applications like robotics, where off-policy algorithms based on Q-function estimation can often attain better sample complexity than direct policy optimization. Existing OPE metrics either rely on a model of the environment, or the use of importance sampling (IS) to correct for the data being off-policy. However, for high-dimensional observations, such as images, models of the environment can be difficult to fit and value-based methods can make IS hard to use or even ill-conditioned, especially when dealing with continuous action spaces. In this paper, we focus on the specific case of MDPs with continuous action spaces and sparse binary rewards, which is representative of many important real-world applications. We propose an alternative metric that relies on neither models nor IS, by framing OPE as a positive-unlabeled (PU) classification problem with the Q-function as the decision function. We experimentally show that this metric outperforms baselines on a number of tasks. Most importantly, it can reliably predict the relative performance of different policies in a number of generalization scenarios, including the transfer to the real-world of policies trained in simulation for an image-based robotic manipulation task.

Off-Policy Evaluation With Online Adaptation for Robot Exploration in Challenging Environments

Beyond Reward: Offline Preference-guided Policy Optimization

Design from Policies: Conservative Test-Time Adaptation for Offline Policy Optimization

Safe Sim-to-Real Robot Exploration with Constrained Bayesian Optimization

Generalize Robot Learning from Demonstration to Variant Scenarios with Evolutionary Policy Gradient

Off-Policy Evaluation via Off-Policy Classification

Learning Off-policy with Model-based Intrinsic Motivation For Active Online Exploration

OPERA: Automatic Offline Policy Evaluation with Re-weighted Aggregates of Multiple Estimators

FOSP: Fine-tuning Offline Safe Policy through World Models

Off-Policy Evaluation for Human Feedback

Adaptive trajectory-constrained exploration strategy for deep reinforcement learning

Off-Policy Evaluation in Doubly Inhomogeneous Environments

Optimistic Active Exploration of Dynamical Systems

Planning to Go Out-of-Distribution in Offline-to-Online Reinforcement Learning

Learning to explore by reinforcement over high-level options

Careful at Estimation and Bold at Exploration

DOP: Deep Optimistic Planning with Approximate Value Function Evaluation

Multi-Objective Autonomous Exploration on Real-Time Continuous Occupancy Maps

KOI: Accelerating Online Imitation Learning via Hybrid Key-state Guidance

Off-OAB: Off-Policy Policy Gradient Method with Optimal Action-Dependent Baseline

Offline-to-Online Multi-Agent Reinforcement Learning with Offline Value Function Memory and Sequential Exploration