DOI: 10.1177/16878132261469440 ISSN: 1687-8132

Scalable multi-AUV complete area coverage with dynamic beacon survey via attention-based MADDPG

Jingyi Xiong, Weijie Zhang, Weiwei Qin, Wenxin Guo, Xinyao Li, Zerui Zhang

Multi-AUV swarms in maritime ISR missions must achieve complete area coverage while performing angle-constrained, multi-round beacon inspections, which are structurally conflicting objectives. Standard MADDPG struggles with this coupling and scalability because its centralized Critic input grows with agent count, causing unstable training and poor transfer. This paper proposes Attention-MADDPG, a unified framework for scalable multi-AUV coverage and dynamic beacon inspection. A Dual-Mode Decision Architecture combines grid memory with the RL policy, switching between coverage exploration and geometric inspection. A multi-head self-attention Critic learns sparse inter-agent interactions, suppressing irrelevant peer information and reducing communication overhead. A bounded composite Grade function rewards coverage, inspection correctness, alignment, efficiency, and cooperative safety. Experiments show that Attention-MADDPG outperforms MATD3, MADDPG, IDDPG, FDSAC, and FTC + CBF in convergence and task Grade. A policy trained with N  = 3 transfers zero-shot to N  = 8, achieving 100% coverage, completing all 152 beacon inspections within 52.37 min, and recording zero cooperative failures. Tests with N in {5, 10, 15} show stable scalability and about 80% fewer active links than full-topology MADDPG at N  = 15. Robustness tests maintain 100% inspection accuracy across eight angle offsets and unchanged performance under sensor noise and packet loss up to 20%.

More from our Archive