DOI: 10.3390/jimaging12080352 ISSN: 2313-433X

Small-Data Deep Learning for Alzheimer-Spectrum Classification from Structural MRI: A Feasibility Study Using OASIS

Ian D. Li, Choong-Yong Ung, Cristina Correia

Accurate estimation of Alzheimer’s disease (AD) severity from structural magnetic resonance imaging (MRI) remains difficult, as disease-associated anatomical alterations are often subtle and publicly available datasets are typically too small to support robust deep learning model training. This feasibility study sought to determine how much Alzheimer’s disease spectrum-related information could be extracted from a small structural MRI cohort using a deliberately lightweight two-dimensional convolutional neural network (2D CNN), and whether transfer learning improves model performance. This study was intended as a methodological proof of concept rather than the development of a clinically deployable diagnostic tool. Structural scans and Clinical Dementia Rating (CDR) labels from the OASIS-1 dataset were filtered to 214 subjects: 124 cognitively normal (CN), 65 with mild cognitive impairment (MCI; CDR = 0.5), and 25 with AD-level impairment (CDR ≥ 1). A compact 2D CNN trained from scratch and a transfer learning model (frozen ImageNet MobileNetV2 features) were evaluated on four binary tasks (CN vs. AD, MCI vs. AD, CN vs. MCI, and CN vs. any impairment) under identical pre-processing and subject-level repeated 5-fold cross-validation (10 repeats), with the decision threshold tuned only on an inner split. Discrimination was summarized by ROC-AUC with 95% confidence intervals (CIs), permutation tests against chance, and per-task sensitivity and specificity. The from-scratch CNN recovered only a broad normal-versus-impaired signal (CN vs. any impairment AUC 0.59) and was at chance on adjacent-stage tasks (MCI vs. AD 0.41; CN vs. MCI 0.51). Transfer learning improved every task: CN vs. AD AUC 0.745 (95% CI 0.730–0.763), CN vs. any impairment 0.642, CN vs. MCI 0.601, and MCI vs. AD 0.599. On an independent OASIS-2 cohort, the transfer learning CN vs. AD model retained AUC 0.748. In this small-data regime, transfer learning recovers substantially more Alzheimer-spectrum signals than a from-scratch CNN, but performance remains modest because it is bounded by CDR-based, non-biomarker-confirmed labels, suggesting the model separates CDR-defined cognitive-status groups rather than detecting AD pathology.

More from our Archive