Abstract:In this paper, we propose a distributed zeroth-order policy optimization method for Multi-Agent Reinforcement Learning (MARL). Existing MARL algorithms often assume that every agent can observe the states and actions of all the other agents in the network. This can be impractical in large-scale problems, where sharing the state and action information with multi-hop neighbors may incur significant communication overhead. The advantage of the proposed zeroth-order policy optimization method is that it allows the agents to compute the local policy gradients needed to update their local policy functions using local estimates of the global accumulated rewards that depend on partial state and action information only and can be obtained using consensus. Specifically, to calculate the local policy gradients, we develop a new distributed zeroth-order policy gradient estimator that relies on one-point residual-feedback which, compared to existing zeroth-order estimators that also rely on one-point feedback, significantly reduces the variance of the policy gradient estimates improving, in this way, the learning performance. We show that the proposed distributed zeroth-order policy optimization method with constant stepsize converges to the neighborhood of a policy that is a stationary point of the global objective function. The size of this neighborhood depends on the agents' learning rates, the exploration parameters, and the number of consensus steps used to calculate the local estimates of the global accumulated rewards. Moreover, we provide numerical experiments that demonstrate that our new zeroth-order policy gradient estimator is more sample-efficient compared to other existing one-point estimators.

Decentralized Natural Policy Gradient with Variance Reduction for Collaborative Multi-Agent Reinforcement Learning

Improved Communication Efficiency in Federated Natural Policy Gradient via ADMM-based Gradient Updates

Global Convergence of Natural Policy Gradient with Hessian-Aided Momentum Variance Reduction

Fast Global Convergence of Natural Policy Gradient Methods with Entropy Regularization.

A Collaborative Multiagent Reinforcement Learning Method Based on Policy Gradient Potential

Provably Fast Convergence of Independent Natural Policy Gradient for Markov Potential Games

Federated Natural Policy Gradient and Actor Critic Methods for Multi-task Reinforcement Learning

Settling the Variance of Multi-Agent Policy Gradients

Scalable and Sample Efficient Distributed Policy Gradient Algorithms in Multi-Agent Networked Systems

Almost Sure Convergence of Networked Policy Gradient over Time-Varying Networks in Markov Potential Games

Augmented Distributed Gradient Methods for Multi-Agent Optimization under Uncoordinated Constant Stepsizes

The Gradient Convergence Bound of Federated Multi-Agent Reinforcement Learning with Efficient Communication.

An off-policy multi-agent stochastic policy gradient algorithm for cooperative continuous control

Cooperative Multi-Agent Reinforcement Learning with Partial Observations

Cooperative Multi-Agent Policy Gradients with Sub-optimal Demonstration

A Policy Gradient Algorithm to Alleviate the Multi-Agent Value Overestimation Problem in Complex Environments

Natural Policy Gradient Primal-Dual Method for Constrained Markov Decision Processes

Decentralized Deep Learning using Momentum-Accelerated Consensus

Multi-agent cooperation through learning-aware policy gradients

Model-Based Decentralized Policy Optimization