Abstract:While decentralized training is attractive in multi-agent reinforcement learning (MARL) for its excellent scalability and robustness, its inherent coordination challenges in collaborative tasks result in numerous interactions for agents to learn good policies. To alleviate this problem, action advising methods make experienced agents share their knowledge about what to do, while less experienced agents strictly follow the received advice. However, this method of sharing and utilizing knowledge may hinder the team's exploration of better states, as agents can be unduly influenced by suboptimal or even adverse advice, especially in the early stages of learning. Inspired by the fact that humans can learn not only from the success but also from the failure of others, this paper proposes a novel knowledge sharing framework called Cautiously-Optimistic kNowledge Sharing (CONS). CONS enables each agent to share both positive and negative knowledge and cautiously assimilate knowledge from others, thereby enhancing the efficiency of early-stage exploration and the agents' robustness to adverse advice. Moreover, considering the continuous improvement of policies, agents value negative knowledge more in the early stages of learning and shift their focus to positive knowledge in the later stages. Our framework can be easily integrated into existing Q-learning based methods without introducing additional training costs. We evaluate CONS in several challenging multi-agent tasks and find it excels in environments where optimal behavioral patterns are difficult to discover, surpassing the baselines in terms of convergence rate and final performance.

Optimistic Exploration Based on Categorical-DQN for Cooperative Markov Games.

Shapley Q-Value: A Local Reward Approach to Solve Global Reward Games

Optimal Exploration Algorithm of Multi-Agent Reinforcement Learning Methods (Student Abstract)

Toward Risk-based Optimistic Exploration for Cooperative Multi-Agent Reinforcement Learning

Optimistic sequential multi-agent reinforcement learning with motivational communication

POAQL: A Partially Observable Altruistic Q-Learning Method for Cooperative Multi-Agent Reinforcement Learning

Exploring Dominant Strategies in Iterated and Evolutionary Games: a Multi-Agent Reinforcement Learning Approach

A Cooperative Multi-Agent Reinforcement Learning Method Based on Coordination Degree

Learning Multi-Agent Cooperation via Considering Actions of Teammates

Learning Automata-Based Multiagent Reinforcement Learning for Optimization of Cooperative Tasks

SCC-rFMQ: a Multiagent Reinforcement Learning Method in Cooperative Markov Games with Continuous Actions

Optimistic Value Instructors for Cooperative Multi-Agent Reinforcement Learning

Multi-robot Cooperation Strategy in a Partially Observable Markov Game Using Enhanced Deep Deterministic Policy Gradient

Multi-Agent Evolutionary Reinforcement Learning Based on Cooperative Games

Lenient Multi-Agent Deep Reinforcement Learning

Towards Cooperation in Sequential Prisoner's Dilemmas: a Deep Multiagent Reinforcement Learning Approach

Self-Motivated Multi-Agent Exploration

Cautiously-Optimistic Knowledge Sharing for Cooperative Multi-Agent Reinforcement Learning

Two Heads Are Better Than One: A Simple Exploration Framework for Efficient Multi-Agent Reinforcement Learning.

Multi-Agent Reinforcement Learning Via Directed Exploration Method

A new multi-agent reinforcement learning approach