Acoustic Features of Sustained Phonation for Schizophrenia Classification: A Feasibility Study
Luka Jelić, Kristian Jambrošić, Vinko Lešić, Pavo OrepićVoice and speech are increasingly studied as indicators of mental health, but the acoustic features of sustained phonation in schizophrenia remain underexplored. This feasibility study examined whether a 500 ms sustained vowel /a/ contains meaningful information for distinguishing patients with schizophrenia from healthy controls. Recordings from 84 participants (41 patients, 43 controls) were analyzed using features from eight acoustic domains. Random Forest, XGBoost, and logistic regression were evaluated under four train–test configurations with data augmentation and participant-grouped cross-validation. Feature importance with Random Forest Gini index, supported by SHAP analysis, produced a compact data-driven 15-feature set (DD15). Random Forest trained with fixed DD15 on original and augmented recordings, and evaluated on original recordings, achieved an AUC of 0.887 ± 0.078, while fold-wise feature selection produced a more conservative 15-feature estimate of 0.837 ± 0.086. Amplitude skewness, harmonic-to-noise ratio, and shimmer measures were the most discriminative features. The conservative DD15 estimate was broadly comparable with the stronger alternative representations while retaining interpretability. Additional sensitivity analyses showed that adjustment for measured recording condition variables reduced, but did not eliminate, internal patient-control discrimination. These findings support the feasibility of sustained-vowel acoustic analysis for schizophrenia classification and emphasize the need for standardized, balanced multi-site validation.