Abstract:Real-world cooperation often requires intensive coordination among agents simultaneously. This task has been extensively studied within the framework of cooperative multi-agent reinforcement learning (MARL), and value decomposition methods are among those cutting-edge solutions. However, traditional methods that learn the value function as a monotonic mixing of per-agent utilities cannot solve the tasks with non-monotonic returns. This hinders their application in generic scenarios. Recent methods tackle this problem from the perspective of implicit credit assignment by learning value functions with complete expressiveness or using additional structures to improve cooperation. However, they are either difficult to learn due to large joint action spaces or insufficient to capture the complicated interactions among agents which are essential to solving tasks with non-monotonic returns. To address these problems, we propose a novel explicit credit assignment method to address the non-monotonic problem. Our method, Adaptive Value decomposition with Greedy Marginal contribution (AVGM), is based on an adaptive value decomposition that learns the cooperative value of a group of dynamically changing agents. We first illustrate that the proposed value decomposition can consider the complicated interactions among agents and is feasible to learn in large-scale scenarios. Then, our method uses a greedy marginal contribution computed from the value decomposition as an individual credit to incentivize agents to learn the optimal cooperative policy. We further extend the module with an action encoder to guarantee the linear time complexity for computing the greedy marginal contribution. Experimental results demonstrate that our method achieves significant performance improvements in several non-monotonic domains.

Value Function Transfer for Deep Multi-Agent Reinforcement Learning Based on N-Step Returns.

Learning in Multi-Agent Systems with Sparse Interactions by Knowledge Transfer and Game Abstraction

Modular Deep Q Networks for Sim-to-real Transfer of Visuo-motor Policies

Cooperative multi-agent target searching: a deep reinforcement learning approach based on parallel hindsight experience replay

Selective Policy Transfer in Multi-Agent Systems with Sparse Interactions

Mutual Information Based Knowledge Transfer Under State-Action Dimension Mismatch

Parallel Knowledge Transfer in Multi-Agent Reinforcement Learning

Adaptive Value Decomposition with Greedy Marginal Contribution Computation for Cooperative Multi-Agent Reinforcement Learning

Enabling Inter-Agent Transfer for Multi-Agent Learning System by Incorporating Role Reversal

Reward-Reinforced Reinforcement Learning for Multi-agent Systems

Value-Decomposition Networks For Cooperative Multi-Agent Learning

Imagined Value Gradients: Model-Based Policy Optimization with Transferable Latent Dynamics Models

DGTRL: Deep graph transfer reinforcement learning method based on fusion of knowledge and data

A Transfer Approach Using Graph Neural Networks in Deep Reinforcement Learning

AVD-Net: Attention Value Decomposition Network for Deep Multi-Agent Reinforcement Learning

Learning to Transfer Role Assignment Across Team Sizes

Transfer in Deep Reinforcement Learning Using Successor Features and Generalised Policy Improvement

Towards Understanding Cooperative Multi-Agent Q-Learning with Value Factorization.

DVF:Multi-agent Q-learning with difference value factorization

Learning Multi-Agent Cooperation via Considering Actions of Teammates