Predicting Complicated Appendicitis: What Can Machine Learning Add?
Mustafa Alper Akay, Ayşe Nur Kübra Kılıç, Ozan Can Tatar, Onursal Varlıklı, Gülşen Ekingen YıldızBackground/Objectives: Early identification of complicated appendicitis in children remains challenging. We developed and internally validated laboratory-based machine-learning models for severity stratification using age and routine admission laboratory data. Methods: This retrospective study included 628 children with surgically confirmed appendicitis treated between 2020 and 2024. Complicated appendicitis was defined by operative or pathological evidence of perforation, gangrene, abscess, phlegmon, diffuse peritonitis, or comparable advanced inflammation. Fifteen candidate predictors were evaluated using five prespecified models. Models were tuned in the training set and evaluated once on an isolated test set. Pairwise DeLong comparisons, decision curve analysis, SHAP, and permutation importance were performed. Results: Complicated appendicitis occurred in 93 patients (14.8%). The prespecified primary CatBoost model achieved a ROC AUC of 0.867, a precision-recall AUC of 0.677, a sensitivity of 0.750, a specificity of 0.863, a positive predictive value of 0.488, and an F1-score of 0.592. Formal comparisons did not demonstrate statistically significant AUC superiority over the other algorithms after Holm correction. Exploratory decision curve analysis showed a greater net benefit than treat-all and treat-none strategies across threshold probabilities of 0.06–0.35. ESR, age, CRP, and CRP-derived indices were the most influential model features. Conclusions: Routine laboratory data may provide adjunctive information for severity stratification, but the modest event count, limited positive predictive value, and absence of external validation preclude stand-alone clinical use.