DOI: 10.3390/drones10100721 ISSN: 2504-446X

Group-Level Heterogeneous Reinforcement-Learning-Guided Jellyfish Search Optimization for Cooperative Multi-UAV Path Planning Under Dynamic and Adversarial Conditions

Nader Alotaibi, Wojdan BinSaeedan

Cooperative multi-UAV path planning must produce coupled trajectories that respect inter-UAV separation constraints through environments with moving obstacles, wind, GPS jamming, and communication loss. Population-based metaheuristics handle the resulting nonconvex, dynamic objectives well, and reinforcement-learning (RL) guidance can adapt their search behavior online; however, existing learned controllers apply one decision homogeneously to the whole swarm, even when part of the fleet is threading a cluttered corridor while the rest cruises in open space. This paper introduces group-level heterogeneous phase control for RL-guided Jellyfish Search Optimization (RL-JSO): UAVs are partitioned into risk-based or random groups, and a Dueling Double DQN selects a joint-phase assignment from an enumerated group-level action space of size 3g, so different groups execute different search phases within the same optimizer iteration. Across three distinct stochastic training replicates and five evaluation campaigns of increasing adversity, group-level control reduces mean mission cost by 21–71% relative to the global RL-JSO controller (Global 3A), raises curriculum mastery from 4.0 to 7.3–8.7 of nine stages, and achieves 95–100% hard-collision-free rates across campaigns under the adopted evaluator, while soft separation margins are frequently violated under the hardest conditions. A risk-versus-random ablation indicates that the specific risk heuristic is not the sole driver of the gains, which are consistent with a substantial role for heterogeneous phase diversity and group-membership dynamics; in deployment-time sensitivity analyses the fixed risk-based policy showed local robustness to moderate risk-weight and group-fraction perturbations, and the overall method ordering was preserved across objective-adaptation strengths. The results support group-level learned optimizer control as a robustness mechanism for cooperative UAV mission planning under the adopted simulation framework.