DOI: 10.3390/jmse14161501 ISSN: 2077-1312

ODARRL: Obstacle- and Disturbance-Aware End-to-End Residual Reinforcement Learning for Underwater Robot Trajectory Tracking with Obstacle Avoidance

Linghan Meng, Zebin Huang, Qingfeng Yao, Yunxiu Zhang, Qifeng Zhang

ROVs are essential for marine exploration and underwater operations, yet conventional teleoperation relies heavily on skilled human operators, and many autonomous methods stop at high-level planning rather than low-level actuation, limiting robustness in disturbed and cluttered environments. This paper proposes ODARRL, an obstacle- and disturbance-aware sensor-to-thruster (ST) end-to-end residual reinforcement learning framework for safe trajectory execution of underwater robots. Using a three-stage curriculum, ODARRL first acquires a basic policy from MPC demonstrations in a static obstacle-free environment, then improves disturbance-robust tracking under random currents, and finally extends to scenarios involving both currents and obstacles. A Dual-Horizon Attention Disturbance Encoder is further designed to capture current-related features from long- and short-term histories, which are fused with robot states and reference information as the input to the ST end-to-end policy. Experiments in Marine Gym with BlueROV2 Heavy demonstrate that ODARRL achieves more stable and robust trajectory tracking under random currents, reducing the mean total tracking error by 69.3%, 31.9%, 45.8%, 73.0% and 25.8% relative to the MPC-imitation policy, PPO, SAC, A2C and VNRS-SAC, respectively. With obstacles introduced, curriculum-initialized policies also exhibit higher path progress and more stable task completion during obstacle-avoidance training.

More from our Archive