Machine Learning Identifies High-Risk Suicide Profiles in a Population-Based Forensic Registry
Alin Ionut Piraianu, Anisia-Luiza Culea-Florescu, Elena Stamate, Ana Fulga, Doriana Iancu, Octavian Stefan Patrascanu, Iuliu FulgaBackground: Suicide is a leading cause of preventable death, yet machine learning (ML) analyses of forensic (medico-legal) suicide data are scarce and, to our knowledge, absent for Romania. Population-based forensic registries offer exhaustive, autopsy-confirmed coverage that is structurally distinct from clinical or civil death-registration data. We applied supervised and unsupervised ML to a complete regional medico-legal suicide registry to profile the method of death and to identify latent victim subgroups of preventive relevance. Methods: We analysed 395 consecutive suicide deaths (Galați and Brăila counties, ≈750,000 inhabitants; 2018–2024). Two supervised classifiers—L2-regularised logistic regression (LR) and random forest (RF, 200 trees)—were trained to discriminate hanging from other methods, using eleven sociodemographic and clinical predictors, and evaluated by 10-fold stratified cross-validation. Given severe class imbalance, the area under the ROC curve (AUC) was the primary metric. Model hyperparameters were fixed a priori, and no class-imbalance correction was applied; both decisions are pre-specified and justified in the Methods. Robustness was assessed by stratified non-parametric bootstrap confidence intervals for the odds ratios, a tipping-point sensitivity analysis for the undocumented clinical fields, and Ward-linkage hierarchical clustering as an independent partitioning check. Predictor importance was quantified by out-of-bag (OOB) permutation importance and Spearman correlations. Unsupervised structure was assessed by K-means clustering (k = 2–7), with the optimal solution selected by the average silhouette coefficient and the elbow (WCSS) criterion. Reporting followed TRIPOD+AI and STROBE. Results: The study population was predominantly male (87.1%) and rural (73.2%), with a mean age of 54.1 years; hanging accounted for 94.2% of deaths—far above the European average (≈50%). RF achieved AUC = 0.865 ± 0.181 and LR AUC = 0.847 ± 0.192, both within the “excellent” discrimination band; sensitivity was very high (0.995–0.997) and specificity was limited (0.233–0.367), an expected consequence of imbalance. Prior suicide attempts (OOB importance 0.959; Spearman ρ = −0.549, p < 0.001; OR = 0.496, 95% CI 0.25–0.72) and the presence of a suicide note (importance 0.575; ρ = −0.439, p < 0.001; OR = 0.549, 95% CI 0.34–0.76) were the dominant predictors. K-means identified two well-separated clusters (silhouette = 0.707), and the partition was reproduced exactly by Ward-linkage hierarchical clustering (adjusted Rand index = 1.000). Cluster 2 (n = 19; 4.8%) was a clinically distinct, younger subgroup (42.4 vs. 54.7 years) characterised by prior attempts (57.9% vs. 0%), suicide notes (68.4% vs. 0%), higher psychiatric comorbidity (52.6% vs. 30.9%) and lower hanging proportion (36.8% vs. 97.1%)—an exploratory, hypothesis-generating profile of recurrent suicidal behaviour with documented prior contact with the medical or medico-legal system. The principal findings were stable across all plausible degrees of clinical under-documentation in the tipping-point sensitivity analysis. Conclusions: ML applied to a complete forensic suicide registry reproduced known regional epidemiology and, beyond classical statistics, isolated an exploratory but clinically coherent high-risk subgroup of direct relevance to the audit of structured post-attempt follow-up. This is, to our knowledge, the first ML study of Romanian forensic suicide data and supports integrating ML into medico-legal research and into the regional targeting and audit of existing post-attempt follow-up provision.