Visual Autonomous Docking for Unmanned Surface Vehicles Using Lightweight Supervised Learning Framework
Junyan He, Wei LiuAutonomous docking is a core capability enabling full autonomy of unmanned surface vehicles (USVs), whose practical deployment demands visual pose estimation with high efficiency, temporal stability, and closed-loop control compatibility. This paper proposes a lightweight monocular visual docking perception framework based on MobileNetV2 and a temporal convolutional network (TCN). In this framework, a MobileNetV2 backbone is adopted to perform end-to-end regression of the USV’s relative pose with respect to the dock from a single monocular image, while a feature-level TCN module fuses sequential visual features across consecutive frames to enhance the short-term stability of pose estimation. To validate the performance and reliability of the proposed method, a high-fidelity simulation environment is established to conduct closed-loop USV docking tests. Comparative results demonstrate that the MobileNetV2 backbone reduces inference latency compared with the VGG19 architecture, and the embedded TCN module effectively suppresses inter-frame pose fluctuations and abnormal estimation jumps. The proposed method provides an efficient and temporally consistent visual perception solution for simulation-validated USV autonomous docking systems.