Abstract:Abstract Intelligent agents and multi-agent systems are increasingly used in complex scenarios, such as controlling groups of drones and non-player characters in video games. In these applications, multi-agent navigation and obstacle avoidance are foundational functions. However, problems become more challenging with the increased complexity of the environment and the dynamic decision-making interactions among agents. The Multi-Agent Deep Deterministic Policy Gradient (MADDPG) algorithm is a classical multi-agent reinforcement learning algorithm successfully used to improve agents’ performance. However, it ignores the temporal message hidden in agents’ interaction with the environment and needs to be more efficient in scenarios with many agents due to its training technique. To address the limitations of MADDPG, we propose to explore modified algorithms of MADDPG for multi-agent navigation and obstacle avoidance. By combining MADDPG with Long Short-Term Memory (LSTM), we obtain the MADDPG-LSTMactor algorithm, which leverages continuous observations over time as input for the policy network, enabling the LSTM layer to capture hidden temporal patterns. Moreover, by simplifying the input of the critic network, we obtain the MADDPG-L algorithm for efficiency improvement in scenarios with many agents. Experimental results demonstrate that these algorithms outperform existing networks in the OpenAI multi-agent particle environment. We also conducted a comparative study of the LSTM-based approach with Transformer and self-attention models in the task of multi-agent navigation and obstacle avoidance. The results reveal that Transformer and self-attention do not consistently outperform LSTM. The LSTM-based model exhibits a favorable tradeoff across varying sequence lengths. Overall, this work addresses the limitations of MADDPG in multi-agent navigation and obstacle avoidance tasks, providing insights for developing intelligent agents and multi-agent systems.

Research on Collision-free Control and Simulation of Single-Agent Based on An Improved DDPG Algorithm

Large Scale Pursuit-Evasion under Collision Avoidance Using Deep Reinforcement Learning.

The Design and Realization of Multi-agent Obstacle Avoidance based on Reinforcement Learning

Path Following for Autonomous Ground Vehicle Using DDPG Algorithm: A Reinforcement Learning Approach

Multi-Agent Path Planning based on MPC and DDPG

Cooperative Path Following Control in Autonomous Vehicles Graphical Games: A Data-Based Off-Policy Learning Approach

Cooperative multi-agent target searching: a deep reinforcement learning approach based on parallel hindsight experience replay

Nobel Lecture. Protein phosphorylation and cellular regulation I.

Bidirectional Obstacle Avoidance Enhancement‐Deep Deterministic Policy Gradient: A Novel Algorithm for Mobile‐Robot Path Planning in Unknown Dynamic Environments

Obstacle Avoidance in Multi-Agent Formation Process Based on Deep Reinforcement Learning

A Collaborative Control Method of Dual-Arm Robots Based on Deep Reinforcement Learning

Mapless Path Planning for Mobile Robot Based on Improved Deep Deterministic Policy Gradient Algorithm

A path planning algorithm fusion of obstacle avoidance and memory functions

Research on Dynamic Path Planning of Mobile Robot Based on Improved DDPG Algorithm

Intelligent Navigation of Indoor Robot Based on Improved DDPG Algorithm

Global Map Assisted Multi-Agent Collision Avoidance Via Deep Reinforcement Learning Around Complex Obstacles

Time-aware MADDPG with LSTM for multi-agent obstacle avoidance: a comparative study

Efficient Multi-agent Navigation with Lightweight DRL Policy

Path planning and dynamic collision avoidance algorithm under COLREGs via deep reinforcement learning

Reciprocal Collision Avoidance for General Nonlinear Agents using Reinforcement Learning