Addressing the Dissimulation Problem in the Assessment of Suicide Risk Using CNN-Based Voice Analysis: A Proof-of-Concept Study
Adela Magdalena Ciobanu, Ana Voichita Tebeanu, Vlad Tau, Anca Ana Chendea, Eduard Dan Franti, Catalin Niculae, Claudiu Ionut Vasile, George Florian Macarie, Liliana Neagu, Monica Dascalu, Cristian Ioan Stoica, Marius Moga, Gabriela IorgulescuBackground/Objectives: Suicide remains a major public health concern, with over 700,000 deaths annually worldwide. Current risk assessment is limited by the tendency of at-risk patients to deny suicidal ideation, with denial rates of approximately 50% among ideators and up to 78% among inpatient suicide decedents. Methods: This proof-of-concept study investigated whether deep learning applied to patient speech could distinguish psychiatric inpatients admitted following a recent suicide attempt from patients with severe depression without documented suicidal ideation. A multi-scale two-dimensional convolutional neural network with residual connections and spatial attention (Multi-Scale CNN) was trained on mel-spectrograms generated exclusively from patient speech extracted by automatic speaker diarization followed by manual verification. The dataset comprised recordings from 88 psychiatric inpatients (59 suicide attempters and 29 patients with severe depression without suicidal ideation). Model performance was evaluated using repeated patient-level stratified five-fold cross-validation, ensuring complete separation of participants between folds. Additional female-only, exploratory male-only, and repeated sex-matched analyses were performed to evaluate the potential influence of sex imbalance. Results: Repeated patient-level five-fold cross-validation yielded a mean classification accuracy of 72.7 ± 6.4% with a mean ROC-AUC of 0.824 ± 0.054, a sensitivity of 88.3 ± 9.5%, and a specificity of 41.3 ± 25.0%. Female-only analysis maintained good discriminative performance (ROC-AUC 0.866 ± 0.123), whereas repeated sex-matched analyses produced lower performance (ROC-AUC 0.621 ± 0.156), indicating that sex imbalance contributed to, but did not fully explain, the observed discrimination. The male-only analysis was considered exploratory because of the limited number of male patients in the depression group. Conclusions: These findings support the feasibility of voice-based patient-level analysis as a complementary objective approach for suicide risk assessment. Although the study remains exploratory because of the limited sample size and the absence of external validation, the revised validation strategy and sex-controlled analyses provide a substantially more robust estimate of model performance. Further multicenter prospective studies are required before clinical implementation.