DOI: 10.3390/brainsci16101027 ISSN: 2076-3425

Machine Learning for Parkinson’s Disease Biomarkers: A Systematic Review of Diagnostic Performance, Risk of Bias, and Clinical Translation Readiness

Osmar Pinto Neto, Tatiana Okubo Rocha Pinho

Background/Objectives: Machine learning (ML) is frequently proposed for Parkinson’s disease (PD) diagnosis, but reported accuracy may not reflect methodological quality, generalizability, or clinical feasibility. We compared diagnostic performance, risk of bias, external validation, and deployment burden across biomarker modalities. Methods: We searched PubMed/Medical Literature Analysis and Retrieval System Online (MEDLINE, Web of Science, and Institute of Electrical and Electronics Engineers (IEEE) Xplore for studies from January 2020 through March 2026 using a Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) 2020-compliant protocol registered with the International Platform of Registered Systematic Review and Meta-analysis Protocols (INPLASY). AI-assisted screening was audited in 147 records (Cohen’s κ = 0.862). Reviewers verified extracted data and Risk Of Bias ASsessment Tool (PROBAST) judgments. Deployment burden was rated across five acquisition domains. Full-text assessment used a modality-stratified sample of retained records; we audited those not selected. Results: Of 8135 deduplicated records, 2046 were retained after screening and 377 reports were sampled for full-text assessment; the final synthesis included 264 studies across nine modalities. Among 252 studies reporting accuracy, median best-model accuracy was 95.3% (range 64.9–100%). However, under our adapted PROBAST rules, 259 studies (98.1%) were at high overall risk of bias; 215 (81.4%) remained at high overall risk when the external-validation item was excluded, and only 20 (7.6%) met the external-validation definition. In ten within-study comparisons, external performance fell in eight. Median deployment burden ranged from 1.0 (voice/speech) to 4.2 (neuroimaging). Conclusions: Reported performance has advanced faster than evidence of generalizability and clinical implementation. Clinical translation will require subject-independent evaluation, external multisite validation, and transparent reporting of performance and deployment requirements. These proportions describe the sample assessed, not the entire PD machine learning literature.