Abstract:While imitation learning requires access to high-quality data, offline reinforcement learning (RL) should, in principle, perform similarly or better with substantially lower data quality by using a value function. However, current results indicate that offline RL often performs worse than imitation learning, and it is often unclear what holds back the performance of offline RL. Motivated by this observation, we aim to understand the bottlenecks in current offline RL algorithms. While poor performance of offline RL is typically attributed to an imperfect value function, we ask: is the main bottleneck of offline RL indeed in learning the value function, or something else? To answer this question, we perform a systematic empirical study of (1) value learning, (2) policy extraction, and (3) policy generalization in offline RL problems, analyzing how these components affect performance. We make two surprising observations. First, we find that the choice of a policy extraction algorithm significantly affects the performance and scalability of offline RL, often more so than the value learning objective. For instance, we show that common value-weighted behavioral cloning objectives (e.g., AWR) do not fully leverage the learned value function, and switching to behavior-constrained policy gradient objectives (e.g., DDPG+BC) often leads to substantial improvements in performance and scalability. Second, we find that a big barrier to improving offline RL performance is often imperfect policy generalization on test-time states out of the support of the training data, rather than policy learning on in-distribution states. We then show that the use of suboptimal but high-coverage data or test-time policy training techniques can address this generalization issue in practice. Specifically, we propose two simple test-time policy improvement methods and show that these methods lead to better performance.

Offline Reinforcement Learning with Value-based Episodic Memory

Model-Based Offline Adaptive Policy Optimization with Episodic Memory

Value Function Evaluation with Data Augmentation for Offline Reinforcement Learning

Value-Evolutionary-Based Reinforcement Learning

Episodic Reinforcement Learning with Expanded State-reward Space

Reward-free Offline Reinforcement Learning

Is Value Learning Really the Main Bottleneck in Offline RL?

Reward-free Offline Reinforcement Learning: Optimizing Behavior Policy Via Action Exploration

Offline RL with No OOD Actions: In-Sample Learning Via Implicit Value Regularization

Sample Efficient Reinforcement Learning Method Via High Efficient Episodic Memory.

Offline Reinforcement Learning With Behavior Value Regularization

Model-based Offline Reinforcement Learning with Lower Expectile Q-Learning

EMaQ: Expected-Max Q-Learning Operator for Simple Yet Effective Offline and Online RL

Real-World Offline Reinforcement Learning from Vision Language Model Feedback

Value-Consistent Representation Learning for Data-Efficient Reinforcement Learning

Pessimistic Nonlinear Least-Squares Value Iteration for Offline Reinforcement Learning

Bi-phase Episodic Memory-Guided Deep Reinforcement Learning for Robot Skills

Weighting Online Decision Transformer with Episodic Memory for Offline-to-Online Reinforcement Learning

Robotic Offline RL from Internet Videos via Value-Function Pre-Training

Offline Reinforcement Learning via High-Fidelity Generative Behavior Modeling

VOCE: Variational Optimization with Conservative Estimation for Offline Safe Reinforcement Learning.