DOI: 10.3390/rs18162791 ISSN: 2072-4292

VCDH-YOLO: Viewpoint-Conditioned Dual-Head Detection for Mixed-View Crack Inspection Across Drone and Ground Platforms

Fangyi Lu, Yifan Hu, Yutong Guo, Zhenglong Ding

Mixed-viewpoint pavement crack detection remains challenging because aerial (Drone) and ground-level (Ground) images exhibit substantially different feature distributions, while repeated downsampling inevitably weakens the representation of fine crack structures in UAV (unmanned aerial vehicles) imagery. To address these issues, this study proposes VCDH-Net, a lightweight mixed-viewpoint crack detection framework built upon YOLOv11n. The framework introduces a Viewpoint-Conditioned Dual-Head Detection (VCDH) architecture that dynamically routes features to viewpoint-specific detection heads through a lightweight viewpoint classifier, enabling specialized optimization while maintaining a shared feature extraction backbone. On this basis, a Lightweight Structure Enhancement (LSE) module is incorporated into mid-level feature layers to reinforce directional crack structures by exploiting local contrast and geometric priors. Furthermore, a Pyramid Detail Refinement (PDR) module is developed for the Drone branch to recover fine-grained spatial information of ultra-small cracks through a lightweight upsample–refine–downsample residual pathway. To provide a more comprehensive evaluation of mixed-viewpoint detection performance, a cross-view assessment framework is further established by introducing three complementary metrics, namely Cross-View Gap (CV-Gap), Worst-View Score (VWS), and Cross-View Balance (CVB). Experiments conducted on a dual-viewpoint pavement crack dataset collected from roads in and around Nanjing, China demonstrate that the proposed method achieves an mAP50-95 of 57.88%, improving the baseline YOLOv11n by 1.55 percentage points. Meanwhile, the Drone-view mAP50-95 increases from 46.98% to 48.83%, and the proposed framework attains the highest CVB score of 0.4965, indicating improved cross-viewpoint detection consistency under the tested conditions. These results, validated on a single-region dataset, demonstrate that VCDH-Net effectively alleviates viewpoint-induced feature discrepancies on the tested data while enhancing the representation of fine crack structures. Generalization to other geographic regions requires further validation on multi-viewpoint datasets not yet publicly available.

More from our Archive