Abstract:One of the main obstacles to broad application of reinforcement learning methods is the parameter sensitivity of our core learning algorithms. In many large-scale applications, online computation and function approximation represent key strategies in scaling up reinforcement learning algorithms. In this setting, we have effective and reasonably well understood algorithms for adapting the learning-rate parameter, online during learning. Such meta-learning approaches can improve robustness of learning and enable specialization to current task, improving learning speed. For temporal-difference learning algorithms which we study here, there is yet another parameter, $\lambda$, that similarly impacts learning speed and stability in practice. Unfortunately, unlike the learning-rate parameter, $\lambda$ parametrizes the objective function that temporal-difference methods optimize. Different choices of $\lambda$ produce different fixed-point solutions, and thus adapting $\lambda$ online and characterizing the optimization is substantially more complex than adapting the learning-rate parameter. There are no meta-learning method for $\lambda$ that can achieve (1) incremental updating, (2) compatibility with function approximation, and (3) maintain stability of learning under both on and off-policy sampling. In this paper we contribute a novel objective function for optimizing $\lambda$ as a function of state rather than time. We derive a new incremental, linear complexity $\lambda$-adaption algorithm that does not require offline batch updating or access to a model of the world, and present a suite of experiments illustrating the practicality of our new algorithm in three different settings. Taken together, our contributions represent a concrete step towards black-box application of temporal-difference learning methods in real world problems.

Off-Policy Training for Truncated TD(\(\lambda \)) Boosted Soft Actor-Critic

Off-Policy Training for Truncated TD(λ) Boosted Soft Actor-Critic.

DSAC-T: Distributional Soft Actor-Critic with Three Refinements

Mitigating Estimation Errors by Twin TD-Regularized Actor and Critic for Deep Reinforcement Learning

TBQ($\sigma$): Improving Efficiency of Trace Utilization for Off-Policy Reinforcement Learning

Modified Retrace for Off-Policy Temporal Difference Learning.

TBQ(σ): Improving Efficiency of Trace Utilization for Off-Policy Reinforcement Learning.

Investigating practical linear temporal difference learning

An Emphatic Approach to the Problem of Off-policy Temporal-Difference Learning

Simplifying Deep Temporal Difference Learning

Revisiting a Design Choice in Gradient Temporal Difference Learning

Context-aware Active Multi-Step Reinforcement Learning

The Ladder in Chaos: A Simple and Effective Improvement to General DRL Algorithms by Policy Path Trimming and Boosting

OPAC: Opportunistic Actor-Critic

META-Learning Eligibility Traces for More Sample Efficient Temporal Difference Learning

A Greedy Approach to Adapting the Trace Parameter for Temporal Difference Learning

Off-Policy Deep Reinforcement Learning Based on Steffensen Value Iteration

Double Actor-Critic with TD Error-Driven Regularization in Reinforcement Learning

Optimizing TD3 for 7-DOF Robotic Arm Grasping: Overcoming Suboptimality with Exploration-Enhanced Contrastive Learning

DoMo-AC: Doubly Multi-step Off-policy Actor-Critic Algorithm