Abstract:One of the main obstacles to broad application of reinforcement learning methods is the parameter sensitivity of our core learning algorithms. In many large-scale applications, online computation and function approximation represent key strategies in scaling up reinforcement learning algorithms. In this setting, we have effective and reasonably well understood algorithms for adapting the learning-rate parameter, online during learning. Such meta-learning approaches can improve robustness of learning and enable specialization to current task, improving learning speed. For temporal-difference learning algorithms which we study here, there is yet another parameter, $\lambda$, that similarly impacts learning speed and stability in practice. Unfortunately, unlike the learning-rate parameter, $\lambda$ parametrizes the objective function that temporal-difference methods optimize. Different choices of $\lambda$ produce different fixed-point solutions, and thus adapting $\lambda$ online and characterizing the optimization is substantially more complex than adapting the learning-rate parameter. There are no meta-learning method for $\lambda$ that can achieve (1) incremental updating, (2) compatibility with function approximation, and (3) maintain stability of learning under both on and off-policy sampling. In this paper we contribute a novel objective function for optimizing $\lambda$ as a function of state rather than time. We derive a new incremental, linear complexity $\lambda$-adaption algorithm that does not require offline batch updating or access to a model of the world, and present a suite of experiments illustrating the practicality of our new algorithm in three different settings. Taken together, our contributions represent a concrete step towards black-box application of temporal-difference learning methods in real world problems.

Metatrace: Online Step-size Tuning by Meta-gradient Descent for Reinforcement Learning Control

Metatrace Actor-Critic: Online Step-Size Tuning by Meta-gradient Descent for Reinforcement Learning Control

Meta-Gradient Search Control: A Method for Improving the Efficiency of Dyna-style Planning

A Greedy Approach to Adapting the Trace Parameter for Temporal Difference Learning

Guarantees for Tuning the Step Size Using a Learning-to-Learn Approach

Adaptive and Multiple Time-scale Eligibility Traces for Online Deep Reinforcement Learning

A Theoretical Understanding of Gradient Bias in Meta-Reinforcement Learning

Train Trajectory Optimization with High-Risk State Space Boundaries: A Safe Reinforcement Learning Approach

Supervised Meta-Reinforcement Learning with Trajectory Optimization for Manipulation Tasks

METRA: Scalable Unsupervised RL with Metric-Aware Abstraction

Learning with Training Wheels: Speeding up Training with a Simple Controller for Deep Reinforcement Learning

No More Pesky Hyperparameters: Offline Hyperparameter Tuning for RL

Adaptive Optimal Control of Nonlinear Systems with Multiple Time-scale Eligibility Traces

Adaptive trajectory-constrained exploration strategy for deep reinforcement learning

META-Learning Eligibility Traces for More Sample Efficient Temporal Difference Learning

Performance-Driven Controller Tuning via Derivative-Free Reinforcement Learning

Meta-Learning for Adaptive Control with Automated Mirror Descent

Meta SAC-Lag: Towards Deployable Safe Reinforcement Learning via MetaGradient-based Hyperparameter Tuning

MESA: Offline Meta-RL for Safe Adaptation and Fault Tolerance

Training a Reinforcement Learning Agent with AutoRL for Traffic Signal Control

On Task-Relevant Loss Functions in Meta-Reinforcement Learning and Online LQR