Abstract:Real-world cooperation often requires intensive coordination among agents simultaneously. This task has been extensively studied within the framework of cooperative multi-agent reinforcement learning (MARL), and value decomposition methods are among those cutting-edge solutions. However, traditional methods that learn the value function as a monotonic mixing of per-agent utilities cannot solve the tasks with non-monotonic returns. This hinders their application in generic scenarios. Recent methods tackle this problem from the perspective of implicit credit assignment by learning value functions with complete expressiveness or using additional structures to improve cooperation. However, they are either difficult to learn due to large joint action spaces or insufficient to capture the complicated interactions among agents which are essential to solving tasks with non-monotonic returns. To address these problems, we propose a novel explicit credit assignment method to address the non-monotonic problem. Our method, Adaptive Value decomposition with Greedy Marginal contribution (AVGM), is based on an adaptive value decomposition that learns the cooperative value of a group of dynamically changing agents. We first illustrate that the proposed value decomposition can consider the complicated interactions among agents and is feasible to learn in large-scale scenarios. Then, our method uses a greedy marginal contribution computed from the value decomposition as an individual credit to incentivize agents to learn the optimal cooperative policy. We further extend the module with an action encoder to guarantee the linear time complexity for computing the greedy marginal contribution. Experimental results demonstrate that our method achieves significant performance improvements in several non-monotonic domains.

Discovering Latent Variables for the Tasks With Confounders in Multi-Agent Reinforcement Learning

Multiagent Reinforcement Learning for Strictly Constrained Tasks Based on Reward Recorder

Optimal Exploration Algorithm of Multi-Agent Reinforcement Learning Methods (Student Abstract)

Modeling the Interaction Between Agents in Cooperative Multi-Agent Reinforcement Learning

LDSA: Learning Dynamic Subtask Assignment in Cooperative Multi-Agent Reinforcement Learning

Multi-Agent Concentrative Coordination with Decentralized Task Representation

Multiagent Continual Coordination via Progressive Task Contextualization

Multi-agent Continual Coordination Via Progressive Task Contextualization

Efficient Multi-Agent Exploration with Mutual-Guided Actor-Critic

Multi-Agent Cooperation via Unsupervised Learning of Joint Intentions

Feudal Latent Space Exploration for Coordinated Multi-Agent Reinforcement Learning.

Two Heads Are Better Than One: A Simple Exploration Framework for Efficient Multi-Agent Reinforcement Learning.

Controlling Large Language Model-based Agents for Large-Scale Decision-Making: An Actor-Critic Approach

Adaptive Value Decomposition with Greedy Marginal Contribution Computation for Cooperative Multi-Agent Reinforcement Learning

Situation-Dependent Causal Influence-Based Cooperative Multi-agent Reinforcement Learning

Credit assignment in heterogeneous multi-agent reinforcement learning for fully cooperative tasks

Episodic Multi-agent Reinforcement Learning with Curiosity-driven Exploration

Imagine, Initialize, and Explore: An Effective Exploration Method in Multi-Agent Reinforcement Learning

LJIR: Learning Joint-Action Intrinsic Reward in cooperative multi-agent reinforcement learning

More Centralized Training, Still Decentralized Execution: Multi-Agent Conditional Policy Factorization