DOI: 10.3390/diagnostics16193139 ISSN: 2075-4418

Diagnostic Performance of Different AI Tools for Chest Radiography: A CT-Anchored Per-Abnormality Analysis in a Tertiary Referral Center

Selin Ardalı Düzgün, İlke Taşçı, Atahan Tamtürk, Nuri Saraç, Melih Karadağ, Mustafa Ege Şeker, Deniz Köksal, Sevinç Sarınç, Meltem Gülsün Akpınar, Figen Başaran Demirkazık, Gamze Durhan

Background/Objectives: Artificial intelligence (AI) increasingly supports chest radiograph interpretation, but per-abnormality comparisons of commercial tools using CT-anchored reference standards remain limited. We compared two commercial AI tools using a CT-anchored, CXR-targeted reference framework. Methods: We retrospectively reviewed chest radiographs obtained between 30 October 2023 and 10 October 2024. Each radiograph was paired with CT performed within 14 days or within 2 days for rapidly progressive findings. Seven abnormalities across five analysis categories were evaluated: fracture, nodule, pleural effusion, pneumothorax, and airspace disease (edema, atelectasis, opacity/consolidation). Radiologists first confirmed abnormalities on CT then assessed their visibility on the paired radiograph. Outputs from vendors A and B were recorded. Sensitivity was calculated for CT-confirmed, radiographically visible abnormalities; specificity, precision, and accuracy were calculated within the same detectability framework. Vendors were compared using McNemar’s test. Results: The cohort included 1707 radiographs from 1680 patients (mean age, 60.3 ± 16.5 years; 50.5% male). Of 3573 CT-confirmed abnormalities, 1981 (55.4%) were radiographically visible and formed the primary reference-positive set. Vendor A had higher sensitivity for fractures, nodules, and pleural effusions (all p < 0.001), whereas vendor B had higher specificity for these findings (all p < 0.05). For pneumothorax, vendor B showed numerically higher sensitivity but lower specificity and accuracy; case numbers were limited, and precision was low for both tools. No significant differences were observed for harmonized airspace disease. Overall specificity was high, while sensitivity and precision varied across abnormalities. Conclusions: Both tools showed abnormality-specific strengths and limitations, supporting tailored implementation and real-world clinical validation.