Better Discrimination, Unchanged Practice: Are Machine Learning Models Ready to Replace Established Risk Scores in Cardiac Surgery? A Narrative Review
Dimitrios E. Magouliotis, Serge Sicouri, Vasiliki Androutsopoulou, Prokopis-Andreas Zotos, Massimo Baudo, Basel RamlawiPreoperative risk stratification underpins consent, treatment selection, and quality benchmarking in cardiac surgery, a task served for two decades by regression-derived scores such as the European System for Cardiac Operative Risk Evaluation II (EuroSCORE II) and the Society of Thoracic Surgeons Predicted Risk of Mortality (STS PROM), together with procedure-specific tools. A rapidly expanding literature reports that machine learning (ML) models achieve higher discrimination than these scores, yet established scores remain the instruments actually used at the bedside. This narrative review examines that paradox. Drawing on studies emphasized between 2023 and 2026, we argue that the reported advantage of ML is real but modest. This advantage is driven primarily by improved discrimination, while key measures of clinical value, including calibration, net benefit, and external or temporal validation, are infrequently reported. We organize the evidence around a four-lens appraisal (discrimination, calibration, clinical utility, and generalizability) and show that most cardiac surgery ML studies focus only on discrimination. We then consider why superior discrimination has not changed practice and outline the evidence needed for an ML-based risk model to justify replacing an established scoring system. The current literature supports a measured conclusion: ML is a discrimination upgrade in search of clinical proof.