Machine Learning Applications in Pediatric Ultrasound for Appendicitis Diagnosis and Severity Stratification: A Systematic Review
Mohammadreza Elhaie, Abolfazl Koozari, Hadis Sharifi, Somayeh Shirazinejad, Qurain Turki Alshammari, Meshari Turki AlshammariABSTRACT
Background
Pediatric acute appendicitis is a leading surgical emergency where ultrasound is the preferred first‐line imaging, yet diagnostic performance varies substantially, prompting interest in machine learning (ML) decision support.
Objective
To synthesize evidence on ML applications using pediatric ultrasound for appendicitis diagnosis and severity stratification.
Methods
Following the Preferred Reporting Items for Systematic Reviews and Meta‐Analyses (PRISMA) 2020 guidelines and the Prediction model Risk of Bias Assessment Tool for Artificial Intelligence (PROBAST‐AI), we prospectively registered the protocol (PROSPERO CRD BLINDED) and searched multiple databases from inception. From 1475 records, 1248 underwent screening after deduplication, 216 full texts were assessed, and 11 studies met inclusion criteria (Cohen's κ = 0.82).
Results
Eleven studies (40–780 patients) were included, with only one cross‐site external validation study. Diagnostic models combining structured ultrasound descriptors with clinical/laboratory features achieved high internal discrimination (AUROC up to 0.993). Across tabular models, internal diagnostic AUROC ranged from 0.91 to 0.96. Severity stratification was less consistent (AUPR 0.70 to AUROC 0.931 internally, dropping to 0.783–0.866 externally). External validation demonstrated transportability limits, with diagnostic AUROC decreasing from 0.96 internally to 0.85 externally. Risk of bias was low in 10/11 studies, though the analysis domain was frequently unclear.
Conclusions
ML‐enhanced pediatric ultrasound demonstrates strong internal diagnostic performance, but severity prediction and cross‐institutional generalizability remain limited, with documented external performance deterioration. Future research should prioritize prospective multi‐institutional evaluation, standardized outcome definitions, calibration reporting, and clinician‐centered interpretability.