Abstract:Recent successes combine reinforcement learning algorithms and deep neural networks, despite reinforcement learning not being widely applied to robotics and real world scenarios. This can be attributed to the fact that current state-of-the-art, end-to-end reinforcement learning approaches still require thousands or millions of data samples to converge to a satisfactory policy and are subject to catastrophic failures during training. Conversely, in real world scenarios and after just a few data samples, humans are able to either provide demonstrations of the task, intervene to prevent catastrophic actions, or simply evaluate if the policy is performing correctly. This research investigates how to integrate these human interaction modalities to the reinforcement learning loop, increasing sample efficiency and enabling real-time reinforcement learning in robotics and real world scenarios. This novel theoretical foundation is called Cycle-of-Learning, a reference to how different human interaction modalities, namely, task demonstration, intervention, and evaluation, are cycled and combined to reinforcement learning algorithms. Results presented in this work show that the reward signal that is learned based upon human interaction accelerates the rate of learning of reinforcement learning algorithms and that learning from a combination of human demonstrations and interventions is faster and more sample efficient when compared to traditional supervised learning algorithms. Finally, Cycle-of-Learning develops an effective transition between policies learned using human demonstrations and interventions to reinforcement learning. The theoretical foundation developed by this research opens new research paths to human-agent teaming scenarios where autonomous agents are able to learn from human teammates and adapt to mission performance metrics in real-time and in real world scenarios.

Convergence of a Human-in-the-Loop Policy-Gradient Algorithm With Eligibility Trace Under Reward, Policy, and Advantage Feedback

Interactive Learning from Policy-Dependent Human Feedback

Human-in-the-Loop Methods for Data-Driven and Reinforcement Learning Systems

Learning Gaussian Policies from Corrective Human Feedback

Interactive Learning with Corrective Feedback for Policies based on Deep Neural Networks

Policy Augmentation: An Exploration Strategy for Faster Convergence of Deep Reinforcement Learning Algorithms

A Single-Loop Deep Actor-Critic Algorithm for Constrained Reinforcement Learning with Provable Convergence

Towards Understanding Asynchronous Advantage Actor-critic: Convergence and Linear Speedup

Counterfactual Multi-Agent Policy Gradients

Maximizing Alignment with Minimal Feedback: Efficiently Learning Rewards for Visuomotor Robot Policy Alignment

Global Convergence of Policy Gradient Methods in Reinforcement Learning, Games and Control

Actor-Critic Reinforcement Learning with Simultaneous Human Control and Feedback

A Minimaximalist Approach to Reinforcement Learning from Human Feedback

The Actor-Advisor: Policy Gradient With Off-Policy Advice

Learning Reward and Policy Jointly from Demonstration and Preference Improves Alignment

Breadcrumbs to the Goal: Goal-Conditioned Exploration from Human-in-the-Loop Feedback

An off-policy multi-agent stochastic policy gradient algorithm for cooperative continuous control

Policy Optimization with Smooth Guidance Learned from State-Only Demonstrations

Human-in-the-Loop Reinforcement Learning in Continuous-Action Space

Policy Gradient from Demonstration and Curiosity

A Multi-Agent Off-Policy Actor-Critic Algorithm for Distributed Reinforcement Learning