Scholay

学术搜索 · AI 审稿 · LaTeX 协作

MASAC-based multi-UAV cooperative detection with beamforming-driven reward design

作者:Yuling Huang, M T Wu, Gong Cheng, Ziwei Wang, Xuanhong Ren, Ying Liu, Qian Chu · 发表于:IET conference proceedings. · 年份:2026 · DOI:10.1049/icp.2026.1415 · 研究领域:UAV Applications and Optimization、Distributed Control Multi-Agent Systems、Reinforcement Learning in Robotics

With their superior maneuverability, economic viability, and robust endurance capabilities, UAVs constitute the fundamental enablers of the emerging low-altitude economic paradigm. When multiple UAVs execute cooperative detection tasks, reinforcement learning is well-suited for decision optimization due to its powerful multi-agent coordination and sequential decision optimization capabilities. However, traditional reward function designs overlook the beamforming principles inherent in detection tasks. This paper proposes a reward function for multi-UAV cooperative detection tasks. Grounded in beamforming principles, an expression for the effective angular range of the main lobe in detection is first derived. This main lobe angle is subsequently incorporated as a key performance metric into the reward function. Furthermore, the complex and dynamic environment of the cooperative detection task is modeled as a Partially Observable Markov Decision Process (POMDP). Based on the above, the Multi-Agent Soft Actor-Critic (MASAC) algorithm is employed to optimize and train the cooperative detection strategies of each UAV. Experimental results demonstrate that UAVs learn optimal cooperative detection coordination strategies through interaction with the environment, achieving a 100% task completion rate and 20.14% formation maintenance rate.