DOI: 10.1097/jp9.0000000000000269 ISSN: 2096-5664

Prediction of organ failure in acute pancreatitis via CT: A multicenter deep learning model with early clinical utility

Yifei Guo, Chengwei Chen, Tiegong Wang, Jiajun Liu, Yixuan Shen, Danqun Zheng, Yilun Zheng, Jieyu Yu, Jing Li, Xu Fang, Fang Liu, Ming Yang, Li Wang, Jianping Lu, Chengwei Shao, Yun Bian

Abstract

Objective:

Acute pancreatitis (AP) incidence is rising globally. Current scoring systems lack sensitivity for early organ failure (OF) prediction and suffer from interobserver variability. This study aimed to develop and validate an artificial intelligence (AI)-driven model for fully automated early prediction of OF in AP using multiphase computed tomography (CT) imaging.

Methods:

This multicenter study included 2746 AP patients from two tertiary hospitals (2011–2024). Patients were split into training ( n =1820), validation ( n =456), and test cohorts ( n =470). An nnMamba-based segmentation model delineated pancreatic/peripancreatic regions on CT. An organ failure risk assessment with CT and learning engine (ORACLE) model integrated deep learning radiomics (severe organ failure–deep learning radiomics [SOF-DLR] score from 57 optimal features) with clinical variables. The primary outcome was OF (Modified Marshall Score≥2).

Results:

OF occurred in 8.7% ( n =240). The ORACLE model achieved the receiver operating characteristic curves (area under the curve [AUCs]) of 0.85 (training), 0.89 (validation), and 0.81 (test), outperforming Modified CT Severity Index (M-CTSI) (AUC 0.68–0.74) and clinical models (AUC 0.67–0.71; DeLong’s P <0.001). The overall negative predictive value for the entire cohort ( n =2746) was 97.2%. High-risk patients ( P >0.700; 1.4% of cohort) had 92.1% OF incidence. The model provided a median early warning time of 3.5 hours (mean 9.17 h) before clinical OF onset, with 55% of cases predicted ≥3 h in advance.

Conclusions:

This AI-based tool enables accurate, automated OF prediction 3.5 h before clinical manifestation, facilitating risk-stratified management. Its generalizability is confirmed in multicenter validation.

More from our Archive