DOI: 10.55994/ejcc.1993408 ISSN: 2667-8721

Beyond the AUROC: data leakage, random splitting, and neglected calibration in critical care machine learning

Ömer Faruk Çakıroğlu
Machine learning is increasingly proposed for predicting triage acuity, deterioration, and mortality in emergency and critical care, yet reported performance is often optimistic. Systematic reviews show that such models are frequently poorly reported and at high risk of bias. We highlight four recurring problems: data leakage, when preprocessing, feature selection, resampling, or hyperparameter tuning extends beyond the training partition; random splitting of a single dataset, which is neither external validation nor statistically efficient; evaluation confined to discrimination or accuracy, disregarding calibration and clinical utility; and unfair comparison of complex algorithms with simpler models and established clinical scores. We urge editors and reviewers to require adherence to current reporting and risk of bias standards, and to constrain claims of clinical readiness without external and prospective validation.

More from our Archive