Beyond Accuracy: A Controlled Comparative Glaucoma Screening Benchmark Across Deep Learning and Hybrid Models Under Within- and Cross-Dataset Conditions
Haifa F. Alhasson, Shuaa S. Alharbi, Muhammed S. AlluwimiBackground/Objectives: Artificial intelligence (AI) models for glaucoma screening using colour fundus photography have shown strong internal performance; however, their external validity and calibration reliability remain uncertain. This study developed a controlled four-dataset benchmark to evaluate six glaucoma-screening models across RIM-ONE, DRISHTI-GS, the Hillel Yaffe Glaucoma Dataset, and ORIGA. Methods: The evaluated models included a hybrid deep-handcrafted random forest (RF), transfer-learning and semi-supervised VGG16 models, and a compact convolutional neural network (CNN). Performance was assessed in within-dataset and cross-dataset settings using discrimination metrics, including accuracy, area under the receiver operating characteristic curve (AUC), balanced accuracy, and Matthews correlation coefficient (MCC), as well as calibration metrics, including Brier score and expected calibration error (ECE). Threshold stability, preprocessing ablation, and repeated-seed analyses were also performed. Results: Within-dataset evaluation showed strong discrimination, with the complete hybrid CNN + histogram of oriented gradients (HOG) + local binary patterns (LBP) + minimum redundancy maximum relevance (mRMR) + RF pipeline achieving a mean accuracy of 0.8697 and an AUC of 0.8972. However, cross-dataset performance was poor, with a mean accuracy of 0.5685 and an AUC of 0.5798. The denoising autoencoder-enhanced transfer-learning model showed improved probability calibration in transfer settings. Preprocessing effects were dataset-dependent, with raw images, region of interest (ROI) cropping, and Retinex normalisation producing different external performance patterns. Conclusions: High internal accuracy did not translate into reliable cross-domain generalisation. None of the evaluated models were suitable for zero-shot deployment without site-specific validation or recalibration.