DOI: 10.3390/bdcc10080270 ISSN: 2504-2289

Evaluating Machine Learning and Deep Learning Models for Early Detection of Alzheimer’s and Parkinson’s Disease: An Explainable Dual-Dataset Study on Clinical Generalizability

Muneera Mohammed Al-Dossary, Atta Rahman

Neurodegenerative diseases such as Alzheimer’s disease (AD) and Parkinson’s disease (PD) pose a significant global healthcare burden due to challenges in early diagnosis. This study investigates the clinical generalizability of machine learning (ML) and deep learning (DL) models for early AD and PD classification across multiple data modalities. A dual-dataset framework was employed, combining global benchmarks (ADNI, PPMI, OASIS, and UCI voice) with local clinical data from King Fahd Hospital of the University (KFHU) in Saudi Arabia. We evaluated ensemble methods, SVMs, neural networks, CNNs, and LSTMs. On structured global data, tree-based ensembles achieved the best performance, with Random Forest reaching 88.24% accuracy for PD and Gradient Boosting achieving 94.44% for AD. For neuroimaging, an LSTM on CNN features attained 98.68% accuracy on a curated MRI dataset. A critical finding was a substantial generalization gap: models excelling on global data showed markedly reduced performance on local KFHU data, with AUC values between 0.50 and 0.77. This degradation is attributed to real-world clinical challenges including severe class imbalance, diagnostic uncertainty in EHRs, and heterogeneous feature representations. The results underscore that data quality and modality are often more consequential than algorithmic complexity. This study provides a reproducible validation framework, highlights the necessity of institution-specific evaluation, establishes performance benchmarks for Saudi healthcare (with the caveat that local sample sizes remain small and the results should be interpreted as exploratory), and demonstrates the use of explainable AI to validate clinical relevance.

More from our Archive