DOI: 10.3390/drones10080600 ISSN: 2504-446X

Spatio-Temporal Attention-Based Improved MADDPG Algorithm for Multi-UAV Formation Path Planning

Dong Zhao, Huaizhi Dong, Wenjing Ren

With the increasing deployment of multi-unmanned aerial vehicle (multi-UAV) systems in dynamic environments, the problem of efficient cooperative path planning has emerged as a critical challenge requiring urgent solutions. To address this issue, this paper proposes a novel joint optimization framework, named spatio-temporal attention-based multi-agent deep deterministic policy gradient (STA-MADDPG). Rather than proposing a new reinforcement learning algorithm in the strict sense, this work integrates advanced spatial-temporal feature extraction with heuristic gradient guidance. First, a cascaded architecture combining multi-head attention and Long Short-Term Memory (LSTM) networks is utilized to extract key local and temporal features, thereby mitigating the dimensionality curse in dense multi-agent observations. Second, an improved dynamic artificial potential field (DAPF) is integrated into the reinforcement learning framework as a state augmentation mechanism, providing heuristic guidance vectors that accelerate convergence and improve obstacle avoidance. Furthermore, to balance computational complexity and adaptive behavior, a rule-based hierarchical formation strategy is designed. The framework maps predefined formations (elliptical, chain, or wedge) to specific environment categories, while the underlying MARL policy governs the dynamic trajectory planning and topology maintenance. Finally, rigorous comparative and ablation experiments are conducted to evaluate path length, search time, and relative position errors. Statistical analysis demonstrates the effectiveness of the proposed framework, achieving up to a 67.3% reduction in search time and a 91.56% search success rate compared with standard MARL baselines in complex environments.

More from our Archive