Machine-Learning Versus Traditional Scores for Predicting Outcomes After Coronary Artery Bypass Graft Surgery: A Systematic Review and Meta-Analysis
Aashray K. Gupta, Ammar Zaka, Daksh Tyagi, Daud Mutahar, Razeen Parvez, Benjamin Muston, Maria Farag, Alexander Lombardo, Aditya Eranki, Ashley Wilson-Smith, Gihwan Song, Brandon Stretton, Joshua G. Kovoor, Stephen Bacchi, Fabio Ramponi, Justin C. Y. Chan, Sarah Zaman, Clara Chow, Pramesh Kovoor, Jayme S. Bennetts, Guy J. MaddernBackground
Coronary artery bypass grafting (CABG) is associated with significant morbidity and mortality. Traditional risk scores, such as the Society of Thoracic Surgery (STS) and EuroSCORE II, have limitations in predicting outcomes, particularly in high-risk patients. Machine learning (ML) models may address these issues by detecting nuanced data patterns not captured by conventional methods. This systematic review and meta-analysis compared the efficacy of ML models with traditional risk scores in predicting outcomes after CABG.
Methods
A comprehensive literature search of records up to August 14 th 2025, was conducted using PubMed, Embase, Web of Science, and the Cochrane Library. Studies included used ML algorithms and traditional risk scores to predict all-cause mortality (in-hospital, 30-day, or longer term as reported by each study) following CABG. Data extraction and quality assessment were independently performed by two reviewers. Meta-analyses were conducted using a linear mixed-effects model, with C-statistics as the primary measure of discrimination.
Results
Twenty-six studies, comprising 565 063 participants, met the inclusion criteria. The pooled C-statistic for ML models was 0.82 (95% CI 0.79-0.85), significantly higher than the 0.73 (95% CI 0.71-0.76) for traditional risk scores (
Conclusions
In this meta-analysis of predominantly internally-validated models, ML approaches showed higher pooled discrimination than traditional risk scores for mortality after CABG. However, given the small number of pooled studies, very high between-study heterogeneity (I 2 = 98%), the predominance of high risk-of-bias studies, and the scarcity of external validation, these findings should be interpreted as supporting the promise of ML rather than establishing proof of clinical superiority. Confirmatory prospective, externally validated studies are required before ML can be recommended for routine pre-operative risk stratification.