DOI: 10.3390/app16189337 ISSN: 2076-3417

Deep Learning for Surface-Conditioned Structural Crack Recognition: Integrating CNN-Attention Models and Calibrated Confidence Estimation

Bonginkosi A. Thango, Sipho G. Thango

Surface recognition can organize structural inspection imagery, but high classification accuracy does not establish crack detection or structural diagnosis. This study evaluates the supplied 49,124-image StructDamage collection, containing nine surface categories and 4586 masonry-focused images. A shared ResNet-18 encoder supports ten convolutional and ten attention-based residual heads under fixed training and validation rules. The validation-selected TokenFormer-8H achieves 99.21% test accuracy and 97.11% macro-F1 on 4926 images. Temperature scaling reduces negative log-likelihood from 0.145 to 0.040 and expected calibration error from 11.18% to 0.33%. However, four selected heads retain zero residual correction, and the macro-F1 improvement over the linear control has a paired 95% component-bootstrap interval spanning zero. A revision audit identifies 101 within-class equal-perceptual-hash groups crossing partitions despite zero exact-hash overlap; excluding affected test components retains 4819 images and 99.19% accuracy. Source-folder labels alone predict 97.97% of test labels, demonstrating strong source–category confounding. A separate ImageNet-only source-held-out probe reaches 83.29% accuracy across 21 eligible folders representing three categories, without establishing nine-class external validity. Additional calibration comparisons, acquisition perturbations, and quantitative attribution tests expose class-dependent uncertainty and sensitivity to image degradation. Blur reduces macro-F1 to 46.81%, and Grad-CAM deletion does not outperform random-pixel deletion. These results support a reproducible surface-category benchmark and an audit of its limitations, not validated crack geometry, mechanism, dimensional measurement, or field safety assessment.