Abstract:This paper is concerned with complex reinforcement learning tasks whose observations are difficult to characterize as appropriate inputs for policy mapping. The representation learning technique is leveraged to extract the features from the observations for optimal action generation. In literature, the action vector, which consists of actions in each dimension, is usually learned from the full features. However, we find empirically that different actions may be only highly related to small part of the features and weakly depend on the rest. Therefore, this unified learning strategy may lead to performance degradation, and a separate learning method is motivated. In this paper, we propose a novel method called Decoupled Reinforcement Learning (DeRL) that decomposes action space by replacing the policy network with decoupled sub-policy group. To cater to all types of tasks where agents. actions in different dimensions can be either weakly correlated or strongly correlated, the Bidirectional Recurrent Neural Network (Bi-RNN) is added as an essential component to capture further shared features for generating more accurate action. In this framework, the decoupled policy network maintains joint representations required by the decision of all actions in different dimensions while decreasing preference in the learning process. In addition, we give a theoretical analysis of DeRL from the perspective of information theory, which shows the difference in information loss between DeRL and others. The performance of the proposed method has been verified by contrastive experiments on 12 tasks, including Mujoco, Atari, and other popular environments.

Hierarchical Reinforcement Learning for Concurrent Discovery of Compound and Composable Policies

Towards Task-Prioritized Policy Composition

Hierarchical Programmatic Reinforcement Learning via Learning to Compose Programs

Developing cooperative policies for multi-stage reinforcement learning tasks

Hierarchical Reinforcement Learning with Opponent Modeling for Distributed Multi-agent Cooperation

Multi-Task Reinforcement Learning in Continuous Control with Successor Feature-Based Concurrent Composition

Planning with a Learned Policy Basis to Optimally Solve Complex Tasks

Bidirectional-Reachable Hierarchical Reinforcement Learning with Mutually Responsive Policies

Data-Efficient Hierarchical Reinforcement Learning for Robotic Assembly Control Applications

DeRL: Coupling Decomposition in Action Space for Reinforcement Learning Task

Hierarchical Reinforcement Learning Based on Planning Operators

Generalization of Compositional Tasks with Logical Specification via Implicit Planning

Hierarchical Reinforcement Learning Based on Continuous Subgoal Space

Cascaded Compositional Residual Learning for Complex Interactive Behaviors

Compositional Policy Learning in Stochastic Control Systems with Formal Guarantees

Hierarchical Subspaces of Policies for Continual Offline Reinforcement Learning

Active Hierarchical Exploration with Stable Subgoal Representation Learning

Continual Task Learning through Adaptive Policy Self-Composition

Policy composition in reinforcement learning via multi-objective policy optimization

Prioritized Soft Q-Decomposition for Lexicographic Reinforcement Learning

Robust Subtask Learning for Compositional Generalization