Abstract:Real-world cooperation often requires intensive coordination among agents simultaneously. This task has been extensively studied within the framework of cooperative multi-agent reinforcement learning (MARL), and value decomposition methods are among those cutting-edge solutions. However, traditional methods that learn the value function as a monotonic mixing of per-agent utilities cannot solve the tasks with non-monotonic returns. This hinders their application in generic scenarios. Recent methods tackle this problem from the perspective of implicit credit assignment by learning value functions with complete expressiveness or using additional structures to improve cooperation. However, they are either difficult to learn due to large joint action spaces or insufficient to capture the complicated interactions among agents which are essential to solving tasks with non-monotonic returns. To address these problems, we propose a novel explicit credit assignment method to address the non-monotonic problem. Our method, Adaptive Value decomposition with Greedy Marginal contribution (AVGM), is based on an adaptive value decomposition that learns the cooperative value of a group of dynamically changing agents. We first illustrate that the proposed value decomposition can consider the complicated interactions among agents and is feasible to learn in large-scale scenarios. Then, our method uses a greedy marginal contribution computed from the value decomposition as an individual credit to incentivize agents to learn the optimal cooperative policy. We further extend the module with an action encoder to guarantee the linear time complexity for computing the greedy marginal contribution. Experimental results demonstrate that our method achieves significant performance improvements in several non-monotonic domains.

Model-based Credit Assignment for Model-free Deep Reinforcement Learning

Learning to assign credit in reinforcement learning by incorporating abstract relations

Deep Reinforcement Learning with Credit Assignment for Combinatorial Optimization

Towards Practical Credit Assignment for Deep Reinforcement Learning

Credit assignment with predictive contribution measurement in multi-agent reinforcement learning

Credit Assignment with Meta-Policy Gradient for Multi-Agent Reinforcement Learning

Shapley Counterfactual Credits for Multi-Agent Reinforcement Learning

Credit assignment in heterogeneous multi-agent reinforcement learning for fully cooperative tasks

Credit Assignment: Challenges and Opportunities in Developing Human-like AI Agents

Asynchronous Credit Assignment Framework for Multi-Agent Reinforcement Learning

Deep reinforcement learning algorithm based on multi-agent parallelism and its application in game environment

Deep reinforcement learning with the confusion-matrix-based dynamic reward function for customer credit scoring

A Survey of Temporal Credit Assignment in Deep Reinforcement Learning

Deep reinforcement learning based on balanced stratified prioritized experience replay for customer credit scoring in peer-to-peer lending

Assigning Credit with Partial Reward Decoupling in Multi-Agent Proximal Policy Optimization

VMAV-C: A Deep Attention-based Reinforcement Learning Algorithm for Model-based Control

MACCA: Offline Multi-agent Reinforcement Learning with Causal Credit Assignment

Credit Assignment During Movement Reinforcement Learning.

Towards Brain-inspired System: Deep Recurrent Reinforcement Learning for Simulated Self-driving Agent

Adaptive Value Decomposition with Greedy Marginal Contribution Computation for Cooperative Multi-Agent Reinforcement Learning

Deep Model-Based Reinforcement Learning for High-Dimensional Problems, a Survey