DOI: 10.3390/diagnostics16162558 ISSN: 2075-4418

Large Language Model Decision Support for Cranial CT in Pediatric Head Trauma

Ezgi Cesur, Ali Halici, Nursel Kurtoglu

Background: Pediatric head trauma is a common reason for emergency department presentation. Although most children have minor injuries, a small proportion harbor clinically important traumatic brain injuries requiring urgent intervention. Artificial intelligence (AI) may offer structured support in computed tomography (CT) decision making, but evidence regarding the performance of general-purpose large language models in pediatric head trauma remains limited. Objective: To evaluate the association between AI-based cranial CT recommendations and clinically meaningful outcomes in pediatric patients with blunt head trauma and to assess the diagnostic performance and clinical utility of the model. Methods: This retrospective single-center observational study included pediatric patients younger than 18 years with blunt head trauma who underwent cranial CT imaging and had complete outcome data. A general-purpose large language model generated binary CT recommendations (“CT recommended” or “CT not recommended”) using structured clinical information available at the time of emergency department presentation. The primary outcome was a composite adverse clinical outcome defined as the occurrence of at least one of the following: emergency surgical intervention, intensive care unit admission, intubation, neurological sequelae or mortality. Diagnostic performance metrics, calibration analysis and decision curve analysis were performed. Results: A total of 819 pediatric patients were included, and the AI model recommended CT in 530 patients (64.7%). The primary outcome occurred in 143 patients (17.5%) and was significantly more frequent in the CT-recommended group than in the CT-not recommended group (24.5% vs. 4.5%; OR 6.90, 95% CI 3.82–12.45; p < 0.001). Abnormal CT findings, emergency surgery, intubation and neurological sequelae were also significantly more common in patients for whom CT was recommended by the AI system. For the primary outcome, the AI recommendation demonstrated a sensitivity of 90.9%, specificity of 40.8%, positive predictive value of 24.5% and negative predictive value of 95.5%. Calibration analysis showed acceptable agreement between predicted probabilities and observed event rates. Decision curve analysis demonstrated greater net benefit than both the “treat-all” and “treat-none” strategies across a range of threshold probabilities. Conclusions: In this clinically selected cohort of pediatric patients with blunt head trauma who underwent cranial CT imaging, AI-based CT recommendations were strongly associated with adverse clinical outcomes and demonstrated high sensitivity and negative predictive value for identifying children at risk of clinically important events. These findings suggest that, within a clinically selected cohort of children who underwent cranial CT imaging, AI-generated CT recommendations were associated with clinically meaningful outcomes. However, these results should not be interpreted as validation of CT decision making in the broader pediatric head trauma population and require prospective validation in unselected cohorts.

More from our Archive