DOI: 10.3390/electronics15153379 ISSN: 2079-9292

Dual-Agent Hierarchical Reinforcement Learning for Typhoon-Avoidance Route Planning of Ships Under Dynamic Wind–Wave–Current Fields

Yu Cai, Ying Li, Liankang Zhang, Jun Song

Ship weather routing in regional seas under severe seasonal weather systems, such as typhoons, poses a critical operational challenge for maritime safety and efficiency. Traditional single-agent reinforcement learning (RL) methods frequently suffer from training instabilities and erratic trajectory adjustments when exposed to the highly non-stationary, multi-scale dynamics of coupled wind–wave–current fields. To address these limitations, this study proposes a dual-agent hierarchical reinforcement learning framework (HRL-MOO) featuring a built-in dynamic risk-weight adaptation mechanism. The proposed global agent discerns large-scale environmental evolutions and adaptively updates the relative weights of wind-, wave-, and current-induced risks at a strategic level, while the local agent translates this macro-level guidance into short-term tactical heading and speed adjustments within realistic vessel motion boundaries. The framework incorporates bathymetric constraints through a high-resolution navigable domain mask derived from ETOPO topography to guarantee practical navigability. Simulated experiments are executed using hourly environmental reanalysis data and best-track records corresponding to the passage of Typhoon Yagi (2024) through the northern South China Sea and Taiwan Strait. The empirical results demonstrate that the cooperative dual-agent structure establishes a superior global Pareto frontier compared to conventional standard single-agent Proximal Policy Optimization (PPO) and metaheuristic baselines. Under extreme typhoon conditions, the HRL-MOO model effectively decreases the cumulative environment-induced risk from approximately 0.17 to 0.12, improves path smoothness by achieving a higher index of 0.842, and accelerates policy convergence to within 3000 training episodes, all while maintaining a highly efficient voyage distance with minimal detour overhead. Ablation studies further validate that the synergy between macro-level stream field recognition and adaptive multi-objective optimization significantly enhances decision-making stability. This framework offers a robust and interpretable computational tool for autonomous ship weather routing in fast-changing, high-risk ocean environments.

More from our Archive