DOI: 10.3390/diagnostics16193190 ISSN: 2075-4418

Machine Learning for Risk Prediction in Colorectal Surgery—Game Changer or Unnecessary? An Institutional Pilot Study

He Ayu Xu, Philip Deslarzes, Jean Louis Raisaro, Dieter Hahnloser, Martin Hübner, Fabian Grass

Background: The aim of the present study was to identify machine learning (ML)-driven risk constellations to predict specific postoperative outcomes and recovery targets to support the development of targeted interventions, optimized to individual risk profiles. Methods: This retrospective institutional cohort study included consecutive patients who underwent elective colorectal surgery over an 11-year study period focusing on baseline patient characteristics, intraoperative details, and postoperative complications. Five supervised classification algorithms for specific complications and recovery targets were used: logistic regression, support vector machine (SVM), XGBoost, decision tree, and gradient boosting. The dataset was randomly divided into training (80%) and testing (20%) subsets. Model performance was assessed using F1-score for classification task, and for each prediction target; the best-performing model and hyperparameter combination were selected based on validation results. To interpret the contribution of individual features to model predictions, SHAP (Shapley Additive exPlanations) analysis was applied to the best-performing model identified for each prediction target. Results: Logistic regression was the best-performing model for six of the seven postoperative outcomes, whereas SVM performed best for postoperative ileus, indicating that optimal model selection was outcome-specific. Overall, however, none of the tested models clearly outperformed the others in their predictive potential. The discriminative power of the models, as measured by ROC-AUC (receiver operating characteristics–area under the curve), was consistently higher than the baseline of 0.50, peaking at 0.93 for renal dysfunction. In terms of classification accuracy, the prediction of postoperative ileus was the most robust, yielding an F1-score of 0.87. For mobilization, the model showed balanced performance across F1-score (0.72), precision (0.74), and recall (0.71). Given the limited performance across all selected models, SHAP analysis was used to improve the interpretability of the predictions to better understand the contributions of each clinical variable. Conclusions: Complex models in the setting of a structured perioperative dataset only slightly outperform simpler approaches, highlighting the importance of data quality and feature richness rather than just increased model complexity. Future work on improving ML-based risk prediction strategies should thus ensure data quality and focus on integrating dynamic perioperative and postoperative measurements, exploring more flexible modeling frameworks, and validating model performance in larger external cohorts.