Diagnostic accuracy of artificial intelligence for fracture detection on radiography and computed tomography: A systematic review and meta-analysis of externally validated clinician-comparative studies
Oscar Luis Castro Guerrero, Jesús Yesith Goenaga Fruto, Laurens Natan Antonio Vásquez, Maria Camila Madariaga, Gloria Ibis Tirado RomeroObjectives:
This systematic review evaluated externally validated, clinician-comparative studies of artificial intelligence (AI)-based fracture detection on conventional radiography and computed tomography (CT) and summarized diagnostic performance by imaging modality.
Material and Methods:
We conducted a systematic review of diagnostic accuracy studies in PubMed/MEDLINE, Embase, and Latin American and Caribbean Health Sciences Literature without language or date restrictions. Eligible studies evaluated AI tools applied directly to radiographs or CT images for fracture detection in living human participants. The focused synthesis included studies with external validation and formal human-reader comparison. Risk of bias and applicability were assessed using the Quality Assessment of Diagnostic Accuracy Studies 2 Tool. Sensitivity and specificity were summarized descriptively and, when complete 2 × 2 data were available, exploratory univariate random-effects meta-analyses were performed separately for radiography and CT. Certainty of the evidence was assessed using the Grading of Recommendations Assessment, Development and Evaluation approach (GRADE).
Results:
Nineteen studies (13 radiography and 6 CT studies) met the inclusion criteria. Anatomical targets included facial bones, upper and lower extremities, pelvis and hip, ribs, and vertebral fractures. Overall risk of bias was low in 2 studies, unclear in 11, and high in 6, with patient selection being the main concern. Four radiography studies contributed complete 2 × 2 data, yielding a pooled sensitivity of 0.86 (95% confidence interval [CI], 0.71– 0.94) and a specificity of 0.84 (95% CI, 0.77–0.89). Three CT studies contributed complete 2 × 2 data, yielding a pooled sensitivity of 0.93 (95% CI, 0.90–0.95) and a specificity of 0.92 (95% CI, 0.81–0.97). GRADE certainty was low for pooled sensitivity and specificity in both modalities and very low for AI-assisted interpretation versus unaided human readers.
Conclusion:
Externally validated AI systems for fracture detection on radiography and CT showed high diagnostic performance and often performed comparably to human readers. However, certainty remains limited by heterogeneity, small numbers of studies with complete 2 × 2 data, and patient-selection concerns. Current evidence supports AI as an assistive tool, but prospective clinically integrated validation is needed before broad implementation.