Automatic Prediction of Vocal Strain Scores in Singing Voice Using Audio and Electroglottographic Modalities
Yuanyuan Liu, Okko Räsänen, Tero Ikävalko, Tua Hakanpää, Vesa Ronkainen, Anne-Maria LaukkanenPurpose:
This study developed machine learning models to predict perceptual strain scores in the singing voice using audio and electroglottographic (EGG) recordings. The study examined the predictive capability of distinct feature sets extracted from audio and EGG modalities and assessed the contributions of participant metadata (META) and feature selection to model performance.
Method:
Data were split into mutually exclusive train-validation (
Results:
During the training-validation phase, the highest Spearman correlations were ρ = .823 (
Conclusions:
Despite the higher numerical performance of audio-based models, EGG features demonstrated robust predictive potential and superior interpretability within the context of professional singing. The results confirm that combining multimodal data with rigorous feature selection provides a robust framework for the objective assessment of singing voice strain. Direct linkage analysis further verified the tight coupling between glottal and acoustic parameters, grounding the multimodal approach in proven vocal physics.