DOI: 10.3390/s26185954 ISSN: 1424-8220

Awareness-Guided Self-Supervised Visual Odometry via Fusion of Pseudo-Depth and Residual-Flow Cues

Shi Zhou, Zijun Yang, Zhen Li, Xianchang Li, Yuchen Sun, Lifeng Zhang

Self-supervised monocular visual odometry (VO) commonly relies on photometric consistency and rigid-scene assumptions to learn camera motion, making pose estimation vulnerable to unreliable cross-frame observations caused by limited geometric observability and independently moving objects. To address this problem, we propose an awareness-guided self-supervised visual odometry framework that emphasizes reliable cross-frame observations for pose learning. Pseudo-depth is exploited to characterize geometric observability and identify regions that are unreliable for cross-frame reconstruction, while residual flow is employed to capture motion inconsistency introduced by independently moving objects. The resulting observable static regions are incorporated into pose estimation at both the input and loss levels, enabling unreliable observations to be suppressed before feature extraction and during optimization. Qualitative experiments on KITTI and Cityscapes verify the effectiveness of pseudo-depth and residual flow in extracting geometric observability and dynamic-motion information across different driving scenes. Quantitative experiments on the KITTI Odometry dataset show that the proposed method achieves average relative translation and rotation errors of 6.04% and 2.42°/100 m, together with an ATE of 0.88 m and an ARE of 0.20°, demonstrating improved ego-motion and trajectory estimation.