DOI: 10.3390/jmse14161537 ISSN: 2077-1312

A PPO-Based Air-Space Collaborative Monitoring Method for Maritime Search and Rescue

Zhaoyan Liao, Zhiqiang Du, Hongyuan Zeng, Kai Liu

Large-scale maritime activity, persistent shipping incidents, and complex marine environments continue to place substantial demands on maritime search and rescue (MSAR). Current MSAR systems do not fully capitalize on the complementary strengths of unmanned aerial vehicles (UAVs) and satellites for collaborative tracking and rescue support. Existing air-space collaboration technologies suffer from two critical limitations: (1) rigid processes, including fixed task allocation, pre-determined path planning without real-time environmental adaptation, and isolated satellite–UAV decision-making, and (2) long task completion cycles, mainly because many methods are adapted to wide-area, long-duration military tracking scenarios. They therefore provide limited support for the dynamic flexibility required in MSAR. This study proposes a Proximal Policy Optimization (PPO)-based air-space collaborative tracking method for maritime moving targets to address these shortcomings and enhance air-space cooperation in MSAR operations. The core implementation of the method includes: (1) integration of target drift forecasting, satellite orbit prediction, UAV task allocation, and path planning into a unified reinforcement learning framework to reduce isolated single-platform decision-making; (2) the adoption of PPO to generate dynamic and flexible air-space collaborative tracking strategies that adjust satellite observation angles and scanning ranges, as well as UAV altitude, speed, and heading according to real-time target, environmental, and platform states; and (3) the design of a multi-dimensional reward function that balances target proximity, energy efficiency, coverage overlap, and inter-platform cooperation to guide strategy optimization. Simulation experiments include system-feasibility verification, baseline-controller comparison, PPO hyperparameter screening, and cross-scenario evaluation. Under idealized communication and payload-matching assumptions, the method enables coordinated tracking of maritime moving targets in simulated MSAR scenarios. In the standardized evaluation, PPO achieved an 11.9% higher mean evaluation episode return, 11.2% lower aggregate UAV energy consumption, and a 9.92-percentage-point greater endurance margin than DDPG. Hyperparameter screening compared candidate learning rates, discount factors, and training budgets, informing the PPO configuration for the subsequent six-scenario evaluation. Across the six controlled scenarios, rewards stabilized after approximately 1400 steps, while action magnitudes varied among regions. These results indicate that the proposed method has potential to enhance air-space collaborative tracking for MSAR decision support.

More from our Archive