DOI: 10.3390/electronics15194463 ISSN: 2079-9292

Intelligent Road Damage Detection from UAV Imagery Using a Context- and Scale-Aware YOLOv10 Architecture

Rong Li, Jian Liu, Cuizhen Sun, Yao Wang, Ying Tian, Xiaojun Huang

Accurate pavement-distress detection in unmanned aerial vehicle (UAV) imagery remains challenging because defects exhibit weak texture, irregular geometry, scale variation, and visual similarity to markings, shadows, water, and repaired pavement. This study presents YOLO-MCG, an engineering adaptation of YOLOv10 that integrates an adapted Context-Guided (CG) feature-fusion block and Multi-Scale Dilated Attention (MSDA) at the stride-8 P3 node, with fixed summation fusion. A direct audit of UAV-PDD2023 identified 11,158 boxes in 2440 images, including 44.8% small targets. Under the primary parent–frame–disjoint split, YOLO-MCG achieved precision 0.781, recall 0.503, mAP@0.5 0.562, and mAP@0.5:0.95 0.308. Across three independent parent–frame–disjoint splits, the corresponding mean metrics were 0.783±0.010, 0.505±0.011, 0.564±0.004, and 0.309±0.006, respectively. The mAP@0.5 improvement over YOLOv10n was statistically significant (p<0.01). Small-object AP@0.5 was 0.433±0.015, compared with 0.643±0.008 for large objects, while the mean AP of the rare repair and pothole classes was 0.535±0.016. Direct-transfer mAP@0.5 on the India, Japan, and Czech subsets of RDD2022 was 0.418±0.015, 0.392±0.018, and 0.374±0.020, respectively. A same-site, multi-condition pilot covering three altitudes and two pavement types obtained recall ranging from 0.43 to 0.52. Given the current limits of split-level resampling, multi-site validation, and survey-grade geolocation, the evidence supports human-in-the-loop screening rather than autonomous or survey-grade inspection.