CT-Referenced Diagnostic Performance and Complementary Error Patterns of Artificial Intelligence Versus a Single Emergency Physician in Chest Radiograph Interpretation: A Retrospective Paired Diagnostic Accuracy Study
Ömer Damar, Mahmut YamanBackground/Objectives: Artificial intelligence (AI) is increasingly used to support chest radiograph interpretation, but its clinical value depends not only on stand-alone accuracy but also on whether its errors differ from those of clinicians. We evaluated CT-referenced, target-specific diagnostic performance and complementary error patterns of a commercial chest radiograph AI system and a single emergency physician. Methods: In this retrospective, single-center, paired diagnostic accuracy study, 920 consecutive adults who underwent posteroanterior chest radiography and thoracic computed tomography (CT) within 4 h during the same emergency department encounter were included. A commercial AI system (hChestXR version 1) and a single blinded emergency physician independently assessed eight prespecified thoracic findings using CT as the reference standard. Paired differences were estimated with 10,000 patient-level bootstrap resamples and tested with exact McNemar tests with Holm adjustment. Exact-target complementarity and simulated same-target OR/AND rules were secondary exploratory analyses. Results: Across 7360 patient–target assessments, including 1232 CT-positive findings, AI had significantly greater Holm-adjusted sensitivity for pneumothorax (90.8% vs. 50.8%), nodule/mass (63.3% vs. 32.7%), and rib fracture (77.2% vs. 46.5%). AI-only detections were most frequent for pneumothorax (47.7%), nodule/mass (44.9%), and rib fracture (43.6%), whereas physician-only detections were most frequent for atelectasis (24.8%), pulmonary edema (21.0%), and lung opacity (19.7%). In exploratory pooled analyses, sensitivity was 73.0% for AI and 62.0% for the physician. Exploratory simulated same-target OR/AND analyses demonstrated a sensitivity–specificity trade-off across pooled patient–target assessments but did not represent observed AI-assisted clinical performance. Conclusions: The AI system and the single emergency-physician reader in the study showed target-dependent, partly non-overlapping error patterns in this CT-selected emergency cohort. The findings support prospective multireader evaluation of AI as a second-reader tool but do not establish benefit from real-time AI-assisted interpretation.