DIDAF-Depth: Dual-Path Interaction and Dual-Attention Fusion Network for Self-Supervised Nighttime Monocular Depth Estimation
Qing Chen, Chao Wei, Bingmeng Zhu, Qiang Yan, Shengbing Chen, Dongmei ZhouSelf-supervised monocular depth estimation is challenging at night because adverse illumination degrades the visual cues required for correspondence estimation and depth inference. Although appearance compensation and domain transfer can mitigate nighttime visual variations, reliable depth recovery under such conditions still depends on exploiting incomplete local structural cues and uncertain scene context. To address this challenge, we propose the Dual-Path Interaction and Dual-Attention Fusion Network (DIDAF-Depth), which improves nighttime depth recovery through coordinated convolutional neural network (CNN)–Transformer interaction, attentional feature fusion, and structure-preserving reconstruction. Specifically, we design the Transformer-CNN Vertical Interaction Fusion (TC-VIF) encoder to perform bidirectional cross-layer exchange, allowing local structural cues and global scene context to complement and progressively refine one another during feature extraction. We further develop the Dual-Coupled Attentional Fusion Module (DCAFM) to model spatial and channel interdependencies and selectively integrate complementary local and global information into a unified representation for depth decoding. Building on DCAFM’s unified representation, we construct the Edge-aware Densely Cascaded Multi-scale Network (EDCMN) to propagate features across scales, reinforce weak boundaries during upsampling, and preserve structural continuity in predicted depth maps. Experiments on the nighttime subsets of the Oxford RobotCar and nuScenes datasets indicate that DIDAF-Depth provides strong and consistent performance under the adopted evaluation protocols, supporting the effectiveness of the proposed framework.