Abstract: Large sequence model (SM) such as GPT series and BERT has displayed outstanding performance and generalization capabilities on vision, language, and recently reinforcement learning tasks. A natural follow-up question is how to abstract multi-agent decision making into an SM problem and benefit from the prosperous development of SMs. In this paper, we introduce a novel architecture named Multi-Agent Transformer (MAT) that effectively casts cooperative multi-agent reinforcement learning (MARL) into SM problems wherein the task is to map agents' observation sequence to agents' optimal action sequence. Our goal is to build the bridge between MARL and SMs so that the modeling power of modern sequence models can be unleashed for MARL. Central to our MAT is an encoder-decoder architecture which leverages the multi-agent advantage decomposition theorem to transform the joint policy search problem into a sequential decision making process; this renders only linear time complexity for multi-agent problems and, most importantly, endows MAT with monotonic performance improvement guarantee. Unlike prior arts such as Decision Transformer fit only pre-collected offline data, MAT is trained by online trials and errors from the environment in an on-policy fashion. To validate MAT, we conduct extensive experiments on StarCraftII, Multi-Agent MuJoCo, Dexterous Hands Manipulation, and Google Research Football benchmarks. Results demonstrate that MAT achieves superior performance and data efficiency compared to strong baselines including MAPPO and HAPPO. Furthermore, we demonstrate that MAT is an excellent few-short learner on unseen tasks regardless of changes in the number of agents. See our project page at https://sites.google.com/view/multi-agent-transformer.

SMART: Sequential Multi-Agent Reinforcement Learning with Role Assignment Using Transformer

S2rl

Multi-Agent Reinforcement Learning is a Sequence Modeling Problem

From Multi-agent to Multi-robot: A Scalable Training and Evaluation Platform for Multi-robot Reinforcement Learning

Multi-Agent Reinforcement Learning with Selective State-Space Models

Efficient Multi-agent Reinforcement Learning by Planning

SC-MAIRL: Semi-Centralized Multi-Agent Imitation Reinforcement Learning

RoMAT: Role-based multi-agent transformer for generalizable heterogeneous cooperation

Sample-efficient multi-agent reinforcement learning with masked reconstruction

Safe Multi-agent Reinforcement Learning with Natural Language Constraints

ACE: Cooperative Multi-agent Q-learning with Bidirectional Action-Dependency

Enabling Multi-Agent Transfer Reinforcement Learning via Scenario Independent Representation

A Multiagent Cooperative Learning System with Evolution of Social Roles

A Multi-agent Cooperative Learning System with Evolution of Social Roles

A New Approach to Solving SMAC Task: Generating Decision Tree Code from Large Language Models

Decentralized Transformers with Centralized Aggregation are Sample-Efficient Multi-Agent World Models

SMARTS: Scalable Multi-Agent Reinforcement Learning Training School for Autonomous Driving

Towards Efficient Multi-Agent Learning Systems

Hierarchical Deep Multiagent Reinforcement Learning with Temporal Abstraction

RODE: Learning Roles to Decompose Multi-Agent Tasks

SMART: Situationally-Aware Multi-Agent Reinforcement Learning-Based Transmissions