Deep Learning and Explainable Artificial Intelligence for Infrared-Thermography-Based CMT-Status Classification in Dairy Cows: A Comparative Evaluation in Egyptian Farms
Alaa T. Elmaria, Elsayed Badr, Sobhy M. A. Sallam, Marwa F. A. Attia, Sara SwedianEarly and accurate identification of mastitis-associated inflammation is critical to increasing animal comfort, reducing financial losses, and enhancing milk quality. Although infrared thermography (IRT) has emerged as a viable non-invasive screening technique, the relative efficacy of convolutional neural networks (CNNs) and Vision Transformers (ViTs) for automated CMT-status classification from thermal udder pictures is still unknown. This work carefully analyzed seven ImageNet-pretrained deep learning architectures using 976 thermal udder images (708 healthy and 268 mastitic images) from 488 Holstein cows (354 healthy cows, 708 images; 134 mastitic cows, 268 images), including two Vision Transformer models (ViT-B/16 and Swin-Tiny) and five CNN models (ResNet-50, DenseNet-121, EfficientNet-B0, MobileNetV2, Inception-V3). Before training, pictures were preprocessed using contrast-limited adaptive histogram equalization (CLAHE), scaled to 224 × 224 pixels, and divided using cow-oriented grouping intended to reduce animal-level data leakage (this grouping could not be independently verified against a ground-truth cow roster; see Limitations). Each model was developed from start to finish and evaluated using an independent hold-out test set and five-fold animal-level cross-validation. DenseNet-121 had the greatest results on the hold-out test set, with an accuracy of 82.2%, an AUC of 0.922, a sensitivity of 90.0%, and a specificity of 79.2%. Additionally, it achieved the highest cross-validation performance (mean AUC = 0.914 ± 0.019). According to statistical analysis, DenseNet-121 was statistically indistinguishable from Inception-V3 under both tests and from ResNet-50 under the more conservative corrected resampled t-test, but significantly outperformed EfficientNet-B0, MobileNetV2, and both Vision Transformers (p < 0.05), while all CNN models outperformed both Vision Transformer models (p < 0.05). Explainable artificial intelligence (XAI) investigation utilizing Grad-CAM corroborated the biological plausibility of the learnt plausibility characteristics by showing that heat patterns in the udder region had a significant impact on model predictions. Grad-CAM analysis showed that heat patterns in the udder region partly drove model predictions, though attention was occasionally influenced by background regions. The best-performing model was tested without retraining on an independent external cohort of 85 cows from a different farm. Although external performance decreased (accuracy = 60.0%; Cohen’s κ = 0.199), consistent with the expected effects of domain shift, the somatic cell count was significantly higher in CMT-positive cows (p < 0.001; AUC = 0.783), supporting the biological significance of the identified thermal abnormalities. Fine-tuned CNNs, led by DenseNet-121, delivered accurate and reproducible thermal CMT-status classification under internal cross-validation, but transportability to an independent farm was limited (external accuracy = 0.600, κ = 0.199), indicating that domain adaptation and multi-farm validation are needed before broader deployment. These findings provide strong support for the application of CNN-based deep learning in IRT-assisted CMT-status screening while highlighting the necessity for larger multicenter datasets to further evaluate transformer-based approaches. aptation and multi-farm validation are needed before broader deployment. These findings provide strong support for the application of CNN-based deep learning in IRT-assisted mastitis screening while highlighting the necessity for larger multicenter datasets to further evaluate transformer-based approaches.