Abstract:Real-world cooperation often requires intensive coordination among agents simultaneously. This task has been extensively studied within the framework of cooperative multi-agent reinforcement learning (MARL), and value decomposition methods are among those cutting-edge solutions. However, traditional methods that learn the value function as a monotonic mixing of per-agent utilities cannot solve the tasks with non-monotonic returns. This hinders their application in generic scenarios. Recent methods tackle this problem from the perspective of implicit credit assignment by learning value functions with complete expressiveness or using additional structures to improve cooperation. However, they are either difficult to learn due to large joint action spaces or insufficient to capture the complicated interactions among agents which are essential to solving tasks with non-monotonic returns. To address these problems, we propose a novel explicit credit assignment method to address the non-monotonic problem. Our method, Adaptive Value decomposition with Greedy Marginal contribution (AVGM), is based on an adaptive value decomposition that learns the cooperative value of a group of dynamically changing agents. We first illustrate that the proposed value decomposition can consider the complicated interactions among agents and is feasible to learn in large-scale scenarios. Then, our method uses a greedy marginal contribution computed from the value decomposition as an individual credit to incentivize agents to learn the optimal cooperative policy. We further extend the module with an action encoder to guarantee the linear time complexity for computing the greedy marginal contribution. Experimental results demonstrate that our method achieves significant performance improvements in several non-monotonic domains.

Reinforcement Learning with Value Function Decomposition for Hierarchical Multi-Agent Consensus Control

Multi-Agent Reinforcement Learning Control for Consensus Problems of Uncertain Nonlinear Multi-Agent Systems

A Hierarchical Control Strategy for the Consensus of Networked Systems

Expert demonstrations guide reward decomposition for multi-agent cooperation

Learning Intra-group Cooperation in Multi-agent Systems.

HAVEN: Hierarchical Cooperative Multi-Agent Reinforcement Learning with Dual Coordination Mechanism

Hierarchical Reinforcement Learning with Opponent Modeling for Distributed Multi-agent Cooperation

Federated Control with Hierarchical Multi-Agent Deep Reinforcement Learning

Value-Decomposition Networks For Cooperative Multi-Agent Learning

Hierarchical Consensus-Based Multi-Agent Reinforcement Learning for Multi-Robot Cooperation Tasks

Data-Based Optimal Consensus Control for Multiagent Systems With Policy Gradient Reinforcement Learning

Value-Decomposition Networks For Cooperative Multi-Agent Learning Based On Team Reward

Distributed learning consensus control based on neural networks for heterogeneous nonlinear multiagent systems

Data-driven output consensus for a class of discrete-time multiagent systems by reinforcement learning techniques

Hierarchical Hybrid Control for Scaled Consensus and Its Application to Secondary Control for DC Microgrid

Understanding Value Decomposition Algorithms in Deep Cooperative Multi-Agent Reinforcement Learning

Adaptive Value Decomposition with Greedy Marginal Contribution Computation for Cooperative Multi-Agent Reinforcement Learning

Modeling the Interaction Between Agents in Cooperative Multi-Agent Reinforcement Learning

Subgoal-based Hierarchical Reinforcement Learning for Multi-Agent Collaboration

Distributed consensus for nonlinear multi-agent systems with two-time-scales: A hybrid reinforcement learning consensus algorithm