Scholay

学术搜索 · AI 审稿 · LaTeX 协作

Multi-Agent Reinforcement Learning With Spatial–Temporal Attention for Flocking With Collision Avoidance of a Scalable Fixed-Wing UAV Fleet

作者:Chao Yan, Chang Wang, Han Zhou, Xiaojia Xiang, Xiangke Wang, Lincheng Shen · 发表于:IEEE Transactions on Intelligent Transportation Systems · 年份:2024 · DOI:10.1109/tits.2024.3505929 · 被引用次数:19 · 研究领域:Distributed Control Multi-Agent Systems、Traffic control and management、Robotic Path Planning Algorithms

Flocking with multiple unmanned aerial vehicles (UAVs) offers significant potential for diverse applications due to its enhanced maneuverability, improved efficiency, and increased robustness. Collision avoidance is a critical and challenging issue for distributed flocking control with a UAV fleet, especially in dynamic environments with varying numbers of non-cooperative intruders. However, existing reinforcement learning based methods mainly focus on flocking with collision avoidance tasks with static obstacles and a fixed number of UAVs. In this article, we propose a scalable multi-agent reinforcement learning based method to solve the distributed flocking with collision avoidance problem for a scalable fleet of fixed-wing UAVs in dynamic environments. Specifically, we cast this problem in a decentralized partially observable Markov decision process framework and propose a scalable multi-agent reinforcement learning algorithm called spatial-temporal attention multi-agent actor-critic (STAAC). In this algorithm, we design a spatial-temporal attention based population-invariant network architecture to facilitate the representation learning of dynamic dimensional observations. By integrating the local spatial attention and global temporal attention mechanisms, STAAC is able to adapt to the changes in the scale of UAV fleets and the number of intruders. Finally, we empirically demonstrate the effectiveness, scalability, and adaptability of the proposed approach in numerical si...