Abstract:Reinforcement learning (RL) theory has largely focused on proving minimax sample complexity bounds. These require strategic exploration algorithms that use relatively limited function classes for representing the policy or value function. Our goal is to explain why deep RL algorithms often perform well in practice, despite using random exploration and much more expressive function classes like neural networks. Our work arrives at an explanation by showing that many stochastic MDPs can be solved by performing only a few steps of value iteration on the random policy's Q function and then acting greedily. When this is true, we find that it is possible to separate the exploration and learning components of RL, making it much easier to analyze. We introduce a new RL algorithm, SQIRL, that iteratively learns a near-optimal policy by exploring randomly to collect rollouts and then performing a limited number of steps of fitted-Q iteration over those rollouts. Any regression algorithm that satisfies basic in-distribution generalization properties can be used in SQIRL to efficiently solve common MDPs. This can explain why deep RL works, since it is empirically established that neural networks generalize well in-distribution. Furthermore, SQIRL explains why random exploration works well in practice. We leverage SQIRL to derive instance-dependent sample complexity bounds for RL that are exponential only in an "effective horizon" of lookahead and on the complexity of the class used for function approximation. Empirically, we also find that SQIRL performance strongly correlates with PPO and DQN performance in a variety of stochastic environments, supporting that our theoretical analysis is predictive of practical performance. Our code and data are available at <a class="link-external link-https" href="https://github.com/cassidylaidlaw/effective-horizon" rel="external noopener nofollow">this https URL</a>.

Dealing with uncertainty: Balancing exploration and exploitation in deep recurrent reinforcement learning

Dealing with uncertainty: balancing exploration and exploitation in deep recurrent reinforcement learning

Fundamental Limits of Reinforcement Learning in Environment with Endogeneous and Exogeneous Uncertainty

Disentangling Uncertainty for Safe Social Navigation using Deep Reinforcement Learning

Beyond Optimism: Exploration With Partially Observable Rewards

Reward Uncertainty for Exploration in Preference-based Reinforcement Learning

Exploration in Feature Space for Reinforcement Learning

Tactical Decision-Making in Autonomous Driving by Reinforcement Learning with Uncertainty Estimation

Is Exploration All You Need? Effective Exploration Characteristics for Transfer in Reinforcement Learning

Enabling risk-aware Reinforcement Learning for medical interventions through uncertainty decomposition

Identifying Critical States by the Action-Based Variance of Expected Return

Identify, Estimate and Bound the Uncertainty of Reinforcement Learning for Autonomous Driving

Decision Making for Human-in-the-loop Robotic Agents via Uncertainty-Aware Reinforcement Learning

The Effective Horizon Explains Deep RL Performance in Stochastic Environments

Exploration in Deep Reinforcement Learning: From Single-Agent to Multiagent Domain

Random Latent Exploration for Deep Reinforcement Learning

Depth and nonlinearity induce implicit exploration for RL

Distributional Reinforcement Learning for Efficient Exploration

Bridging the gap between Markowitz planning and deep reinforcement learning

Adaptive trajectory-constrained exploration strategy for deep reinforcement learning

The Uncertainty Bellman Equation and Exploration