Scalable multi-AUV complete area coverage with dynamic beacon survey via attention-based MADDPG
Jingyi Xiong, Weijie Zhang, Weiwei Qin, Wenxin Guo, Xinyao Li, Zerui Zhang
Multi-AUV swarms in maritime ISR missions must achieve complete area coverage while performing angle-constrained, multi-round beacon inspections, which are structurally conflicting objectives. Standard MADDPG struggles with this coupling and scalability because its centralized Critic input grows with agent count, causing unstable training and poor transfer. This paper proposes Attention-MADDPG, a unified framework for scalable multi-AUV coverage and dynamic beacon inspection. A Dual-Mode Decision Architecture combines grid memory with the RL policy, switching between coverage exploration and geometric inspection. A multi-head self-attention Critic learns sparse inter-agent interactions, suppressing irrelevant peer information and reducing communication overhead. A bounded composite Grade function rewards coverage, inspection correctness, alignment, efficiency, and cooperative safety. Experiments show that Attention-MADDPG outperforms MATD3, MADDPG, IDDPG, FDSAC, and FTC + CBF in convergence and task Grade. A policy trained with