DOI: 10.1177/20552076261492122 ISSN: 2055-2076

A multicenter-validated machine learning model for predicting in-hospital mortality in ICU patients with cirrhosis

Qinhua Yang, Zhikun Xu, Zeming Chen, Weifeng Li, Xiao Tang, Yijing Su, Dongting Peng, Boru Wu, Zhiming Chen, Yichun Jiang, Xueyan Liu, Xiqiu Yu

Background

Intensive care unit (ICU) patients with cirrhosis face a high in-hospital mortality risk. We aimed to develop and externally validate an interpretable machine learning model for early mortality prediction using routinely collected variables from the first 24 hours of ICU admission.

Methods

In this multicenter retrospective cohort study, the model was developed and internally validated on the eICU Collaborative Research Database (eICU) and externally validated on Medical Information Mart for Intensive Care-IV (MIMIC-IV), Northwestern ICU (NWICU), and Shenzhen People’s Hospital (SZPH). The primary outcome was all-cause in-hospital mortality. Features selected by LASSO were used to build an XGBoost model. Performance was evaluated by AUROC and Brier score, with calibration curves and decision curve analysis, and compared with five established scores (MELD, MELD-Na, Child-Pugh, SOFA, CLIF-SOFA) and logistic regression using DeLong’s tests. SHAP provides interpretable feature attribution.

Results

We included 3,891 patients with cirrhosis (eICU, n=1,870; MIMIC-IV, n=1,697; NWICU, n=134; SZPH, n=190); in-hospital mortality rates were 19.79%, 28.76%, 38.81%, and 47.37%, respectively. The final model used 10 predictors (prothrombin time, mechanical ventilation, creatinine, mean arterial pressure, lactate, respiratory rate, total bilirubin, oxygen saturation, heart rate, and white blood cell count). It showed good discrimination and calibration internally (AUROC 0.814; Brier 0.125) and generalized to external cohorts (AUROCs 0.781, 0.733, and 0.719 in MIMIC-IV, NWICU, and SZPH, respectively). In MIMIC-IV cohort, XGBoost significantly outperformed MELD-Na, Child-Pugh, and SOFA, while differences versus MELD, CLIF-SOFA, and logistic regression were not significant. In eICU internal validation cohort, XGBoost significantly outperformed Child-Pugh and SOFA but not other comparators; in SZPH cohort, it showed comparable discrimination to all comparators.

Conclusion

We developed and multicenter-validated an interpretable XGBoost model predicting in-hospital mortality in ICU patients with cirrhosis, with competitive or superior discrimination versus established scores. Prospective validation, local recalibration, and implementation studies are warranted.