SDFA-Net: Urban Manhole Cover Detection via Fisheye Image and LiDAR Point Cloud Fusion
Qiuping Lan, Shuwen Hu, Jia Li, Yijie Huang, Zikuan Li, Qiang FanAutomated manhole cover detection, a key task in urban infrastructure inspection, is hindered by modality-specific limitations: conventional monocular image-based detectors do not directly quantify cover-to-road elevation differences, LiDAR-based methods suffer from sparse sampling that limits recall for small, distant targets, and existing fusion approaches suffer from cross-modal spatial misalignment and lack 3D defect quantification. This paper proposes SDFA-Net, an end-to-end fisheye image and LiDAR point cloud fusion framework for manhole cover detection and diagnosis in wearable mobile mapping scenarios. The semi-dense depth-supervised view transformation (SDS-VT) module constructs semi-dense depth ground truth from multi-frame accumulated point clouds to explicitly supervise image-to-bird’s-eye-view (BEV) feature projection, addressing cross-modal spatial registration under wide-angle fisheye distortion. The fidelity-adaptive fusion (FA-Fusion) module builds a physical fidelity map by combining point-cloud projection density with depth-prediction entropy, then performs quality-aware adaptive fusion via cascaded channel-spatial attention. The geometry-aware decoupled detection head (Geo-Head) adopts an anchor-free decoupled architecture to jointly predict 2D localization, physical dimensions, and millimeter-level relative elevation in a single forward pass. A multimodal dataset of 3500 spatiotemporally aligned frames is constructed using the NavVis VLX wearable platform, covering five defect categories, with Random Sample Consensus (RANSAC)-derived relative-elevation reference labels independently validated by digital leveling. Experiments show that SDFA-Net achieves 92.9% mAP@0.5 and 63.5% mAP@0.5:0.95 with 18.5 M parameters, improving over the best single-modal baseline by 4.6, 11.0, and 8.7 percentage points in mAP@0.5, mAP@0.5:0.95, and Recall, respectively. Compared with BEVFusion, SDFA-Net achieves higher detection accuracy while maintaining a smaller overall model size and an inference speed of 35 FPS.