DOI: 10.1177/15533506261468192 ISSN: 1553-3506

Machine-Learning Versus Traditional Scores for Predicting Outcomes After Coronary Artery Bypass Graft Surgery: A Systematic Review and Meta-Analysis

Aashray K. Gupta, Ammar Zaka, Daksh Tyagi, Daud Mutahar, Razeen Parvez, Benjamin Muston, Maria Farag, Alexander Lombardo, Aditya Eranki, Ashley Wilson-Smith, Gihwan Song, Brandon Stretton, Joshua G. Kovoor, Stephen Bacchi, Fabio Ramponi, Justin C. Y. Chan, Sarah Zaman, Clara Chow, Pramesh Kovoor, Jayme S. Bennetts, Guy J. Maddern

Background

Coronary artery bypass grafting (CABG) is associated with significant morbidity and mortality. Traditional risk scores, such as the Society of Thoracic Surgery (STS) and EuroSCORE II, have limitations in predicting outcomes, particularly in high-risk patients. Machine learning (ML) models may address these issues by detecting nuanced data patterns not captured by conventional methods. This systematic review and meta-analysis compared the efficacy of ML models with traditional risk scores in predicting outcomes after CABG.

Methods

A comprehensive literature search of records up to August 14 th 2025, was conducted using PubMed, Embase, Web of Science, and the Cochrane Library. Studies included used ML algorithms and traditional risk scores to predict all-cause mortality (in-hospital, 30-day, or longer term as reported by each study) following CABG. Data extraction and quality assessment were independently performed by two reviewers. Meta-analyses were conducted using a linear mixed-effects model, with C-statistics as the primary measure of discrimination.

Results

Twenty-six studies, comprising 565 063 participants, met the inclusion criteria. The pooled C-statistic for ML models was 0.82 (95% CI 0.79-0.85), significantly higher than the 0.73 (95% CI 0.71-0.76) for traditional risk scores ( P < 0.0001). The top-performing ML model achieved a C-statistic of 0.98 (CI 0.95-1.00). Calibration was reported inconsistently across studies and was synthesised narratively rather than quantitatively. Where reported, ML calibration was generally adequate but a robust head-to-head comparison with traditional risk scores was not possible. Subgroup analyses revealed consistent superior performance of ML models across various algorithms and covariate sets.

Conclusions

In this meta-analysis of predominantly internally-validated models, ML approaches showed higher pooled discrimination than traditional risk scores for mortality after CABG. However, given the small number of pooled studies, very high between-study heterogeneity (I 2 = 98%), the predominance of high risk-of-bias studies, and the scarcity of external validation, these findings should be interpreted as supporting the promise of ML rather than establishing proof of clinical superiority. Confirmatory prospective, externally validated studies are required before ML can be recommended for routine pre-operative risk stratification.

More from our Archive