Artificial intelligence and machine learning for prediction of postoperative complications: A systematic review focused on anesthesiology, perioperative risk stratification, and clinical applicability
Hamza Hafiani, Hamza Kassimi, Salma Bekkour, Rim Bentalba, Khalil Abou ElalaaAbstract
Artificial intelligence (AI) and machine learning (ML) are increasingly used to predict postoperative complications, but their clinical value for anesthesiologists depends on more than discrimination alone. This systematic review evaluated AI- and ML-based models for predicting postoperative complications in surgical patients, with emphasis on model performance, validation, calibration, timing of predictor availability, interpretability, and clinical applicability. The review followed Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) 2020 principles. Eligible studies developed, validated, or compared AI/ML models for postoperative complication prediction and reported at least one performance metric. Risk of bias was assessed using Prediction model Risk of Bias ASsessment Tool (PROBAST), and reporting quality was interpreted in light of Transparent Reporting of a multivariable prediction model for Individual Prognosis Or Diagnosis (TRIPOD) principles. Thirty-five studies were included. The most frequently predicted outcomes were acute kidney injury, postoperative delirium, cardiovascular complications, infectious complications, and composite postoperative morbidity. Reported best-model discrimination ranged approximately from 0.70 to 0.97. Models incorporating intraoperative variables often showed stronger performance than preoperative-only models, particularly for acute kidney injury and infectious complications, but this reduced their usefulness for early preoperative decision-making. External validation was uncommon, calibration was inconsistently reported, and the analysis domain was the most frequent source of concern in PROBAST assessment. Current AI/ML models for postoperative complication prediction are promising, but their routine use in anesthesiology requires stronger external validation, transparent calibration, workflow integration, and clinically interpretable decision-support strategies.