Early Detection and Severity Assessment of Parkinson's Disease Using Speech-Based Machine Learning
Ranjeet Kumar, Akkaraju Venkata Ambareesh, Shaik Mohammad Rabbani, Chatwada Muni Vijayesh Singh, Rekkala Venkateswara ReddyParkinson's disease (PD) is a neurodegenerative disease characterized by progressive motor disability and early speech disorders. Early diagnosis and regular follow-up allow timely clinical intervention and a better quality of life. This paper introduces a machine learning framework that identifies PD and estimates its severity from acoustic features derived from speech. The data analyzed are from the Parkinson's Telemonitoring dataset of the UCI Machine Learning Repository, which consists of 5,875 voice recordings. Acoustic characteristics such as jitter, shimmer, harmonicity, and nonlinear vocal parameters capture the speech changes associated with PD. Supervised linear regression, support vector regression (SVR), random forest (RF) regression, and gradient boosting (GB) models, together with an RF and GB ensemble, were trained to predict the motor and total Unified Parkinson's Disease Rating Scale (UPDRS) scores, and the predicted scores were then grouped into discrete severity levels. Model performance was analyzed with accuracy, precision, recall, and F1 score. The RF model achieved the best accuracy, 95.49%, which indicates strong predictive ability.