MIXRTs: Toward Interpretable Multi-Agent Reinforcement Learning Via Mixing Recurrent Soft Decision Trees

Zichuan Liu,Yuanyang Zhu,Zhi Wang,Yang Gao,Chunlin Chen
DOI: https://doi.org/10.48550/arxiv.2209.07225
2022-01-01
Abstract:While achieving tremendous success in various fields, existing multi-agentreinforcement learning (MARL) with a black-box neural network architecturemakes decisions in an opaque manner that hinders humans from understanding thelearned knowledge and how input observations influence decisions. Instead,existing interpretable approaches, such as traditional linear models anddecision trees, usually suffer from weak expressivity and low accuracy. Toaddress this apparent dichotomy between performance and interpretability, oursolution, MIXing Recurrent soft decision Trees (MIXRTs), is a novelinterpretable architecture that can represent explicit decision processes viathe root-to-leaf path and reflect each agent's contribution to the team.Specifically, we construct a novel soft decision tree to address partialobservability by leveraging the advances in recurrent neural networks, anddemonstrate which features influence the decision-making process through thetree-based model. Then, based on the value decomposition framework, we linearlyassign credit to each agent by explicitly mixing individual action values toestimate the joint action value using only local observations, providing newinsights into how agents cooperate to accomplish the task. Theoretical analysisshows that MIXRTs guarantees the structural constraint on additivity andmonotonicity in the factorization of joint action values. Evaluations on thechallenging Spread and StarCraft II tasks show that MIXRTs achieves competitiveperformance compared to widely investigated methods and delivers morestraightforward explanations of the decision processes. We explore a promisingpath toward developing learning algorithms with both high performance andinterpretability, potentially shedding light on new interpretable paradigms forMARL.
What problem does this paper attempt to address?