Automated Concrete Structural Damage Classification via Transfer Learning Ensembles and Explainable AI: A Multi-Dataset Validation Study
Hernán Patricio Moyano-Ayala, Omar Sebastian Muñoz-Merino, Diego Fernando Mayorga-Pérez, Jorge Sebastián Buñay-GuamánMaintaining the structural integrity of concrete infrastructure is vital for public safety, yet conventional visual inspections remain subjective, costly and difficult to scale. This study reports a controlled benchmark of transfer-learning classifiers for patch-level binary crack detection on concrete surfaces, together with a soft-voting ensemble that combines two complementary convolutional backbones: ResNet-50 for deep hierarchical feature extraction and MobileNetV2 for lightweight inference. The two branches are fused through a single scalar weight α selected on the validation partition. Six transfer-learning baselines are trained under identical splits, seeds and augmentation settings on two public benchmarks: the Özgenel and Sorguç concrete crack image set (40,000 images, balanced classes, hosted on Mendeley Data) and SDNET2018 (56,092 images covering bridge decks, walls and pavements, with an approximate 5.6:1 class imbalance). Each model is trained and tested within a single dataset; no cross-dataset transfer experiment was performed, so the design constitutes multi-dataset validation rather than a demonstration of cross-domain generalisation. The ensemble attains an F1-score of 0.9995 (AUC = 1.0000) on the balanced benchmark and 0.8083 (AUC = 0.9500) on SDNET2018; in both cases, it is marginally above the best individual model. Because a single random seed was used, these differences are reported without significance testing. Crack recall on SDNET2018 is 0.7309, so approximately one cracked surface in four is missed, and the pipeline is accordingly positioned as a human-in-the-loop screening aid rather than an autonomous inspection system. Inference latency is 12.9 ms per image at a total model size of 98.5 MB on a desktop-class graphics processing unit (GPU); no embedded platform was evaluated. Grad-CAM maps are presented as a qualitative visual audit of the individual backbones, not as validated evidence of explanation faithfulness.