From Single-Center Feasibility to a Multi-Institutional Cohort: A 2.5D Deep Learning Study of Pediatric Craniofacial Fracture Triage on CT
Bartosz Ignac, Łukasz Walusiak, Natalia Sitek-Ignac, Katarzyna Tyburska, Tomasz Wach, Damian Dudek, Krzysztof Dowgierd, Marcin Kozakiewicz, Zygmunt Wróbel, Bogusława Orzechowska-WylęgałaBackground: Pediatric craniofacial fractures are difficult to assess on computed tomography (CT) because growth-related anatomy and subtle fracture patterns can obscure clinically relevant findings. We evaluated a lightweight 2.5D deep-learning approach after expansion of an earlier 63-examination pilot. Methods: The retrospective cohort comprised 209 fully anonymized pediatric craniofacial CT examinations from 209 unique patients, with exactly one examination per patient (103 fracture-positive and 106 fracture-negative). The center contributions were Katowice (n = 135), Olsztyn (n = 34), and Lodz (n = 40). Central axial, coronal, and sagittal reconstructions were stacked as a three-channel input to an ImageNet-initialized ResNet-18. The historical experiment retained a stratified 146/63 training/development–validation split. Reviewer-requested post hoc analyses comprised descriptive calibration, controlled retraining without vertical flipping, three repetitions of five-fold cross-validation, input and initialization ablations, and leave-one-center-out testing. Results: On the 63-examination development-validation subset, AUROC was 0.934 (95% CI 0.862–0.989) and AUPRC was 0.938 (95% CI 0.878–0.987). Repeated five-fold cross-validation yielded mean AUROC 0.899 (SD 0.033) and AUPRC 0.891 (SD 0.040). Leave-one-center-out AUROC varied from 0.604 (SD 0.025) for Katowice to 0.793 (SD 0.097) for Lodz; the Olsztyn holdout permitted sensitivity-only assessment. Conclusions: The original split and repeated cross-validation support an internal feasibility signal, but the post hoc center-held-out results show that performance was not uniformly transportable across institutions. The findings do not establish external validity or clinical utility.