DOI: 10.1061/jccee5.cpeng-7232 ISSN: 0887-3801

Enhanced Road Crack Segmentation Using Self-Supervised Learning in Motion-Blurred Images

Ye Liu, Jinyan Feng, Bowen Du, Jun Chen

Abstract

With the increasing deployment of autonomous inspection platforms [e.g., unmanned aerial vehicles (UAVs) and robots], the segmentation performance of learning-based road inspection methods is facing severe challenges due to motion blur. In response, an innovative image segmentation framework is proposed in this study that treats motion deblurring as a pretext task for self-supervised learning. The proposed framework uses a shared encoder to process both motion deblurring and segmentation models. Structural features from the low-level visual task are effectively transferred to the high-level semantic task without introducing additional data. Consequently, the efficiency, performance, and robustness of the segmentation model on motion-blurred images are significantly enhanced. Crucially, this single-model approach eliminates the need for a separate, computationally expensive restoration step, making it particularly efficient for resource-constrained edge devices compared to traditional multistage pipelines. This framework proves adaptable to datasets of varying complexities and network architectures. On a simple motion-blurred crack segmentation dataset captured from UAVs, the CNN-based DeblurGAN enhanced by this framework improved the crack Intersection over Union (IoU) from 49.39% to 70.51%, closely matching the performance achieved using sharp training data (71.98%). Furthermore, on a complex dataset captured from terrestrial vehicles, the proposed framework successfully prevented the “model collapse” observed in the Transformer-based Restormer. The augmented model achieved a mean F1 and mean IoU of 88.19% and 80.65%, respectively, surpassing the next best-performing ConvNext (87.34% and 79.53%). Future studies will explore refining the network architecture to further integrate deblurring and segmentation characteristics and extending this self-supervised paradigm to unified perception frameworks robust against a wider spectrum of real-world visual degradations.

More from our Archive