DOI: 10.1136/bmjopen-2026-124868 ISSN: 2044-6055

Diagnostic accuracy of a DenseNet-121 deep learning algorithm for chest radiograph triage in health assessment applicants: a prospective shadow-mode validation study in Nepal

Lochan Shrestha, Dinesh Maharjan, Uttam Bista

Objectives

To evaluate the diagnostic accuracy of a publicly available DenseNet-121 convolutional neural network (TorchXRayVision) for triaging chest radiographs of health assessment applicants at a tertiary hospital in Nepal.

Design

Prospective, single-centre, shadow-mode diagnostic accuracy validation study. Reported in accordance with the Standards for Reporting of Diagnostic Accuracy Studies (STARD) 2015 checklist and STARD-Artificial Intelligence (AI)/Developmental and Exploratory Clinical Investigations of DEcision support systems driven by Artificial Intelligence (DECIDE-AI) guidelines.

Setting

Department of Radiology and Imaging, Patan Academy of Health Sciences/Patan Hospital, Lalitpur, Nepal.

Participants

826 consecutive health assessment applicants (foreign employment predeparture medical examination and student migration) undergoing chest radiography from 5 June 2026 to 20 June 2026 inclusive (16 days). Two cases were excluded due to Digital Imaging and Communications in Medicine technical failure.

Index test

DenseNet-121 algorithm (TorchXRayVision library, densenet121-res224-all pretrained weights). A maximum aggregated pathology probability score was derived per radiograph and compared against a post hoc derived threshold of 0.6258 (selected as the highest threshold achieving the prespecified ≥95% sensitivity criterion).

Reference standard

Single-reader-per-case review by one of three radiologists—two board-certified radiodiagnosticians (LS: 276 cases; DM: 275 cases) and one radiology resident (UB: 275 cases)—each blinded to AI output, using a standardised data collection worksheet capturing binary classification (abnormal/normal) and free-text findings.

Results

Of 826 radiographs, 41 (4.97%) were classified as abnormal by the reference standard. At the post hoc derived threshold of 0.6258, the DenseNet-121 algorithm achieved: sensitivity 95.12% (95% CI 83.9% to 98.7%), specificity 77.2% (95% CI 74.1% to 80.0%), area under the receiver operating characteristic curve 0.9583 (95% bootstrap CI 0.9225 to 0.9843), negative predictive value (NPV) 99.67% (95% Wilson CI 98.8% to 99.9%), positive predictive value 17.89% (95% Wilson CI 13.4% to 23.5%) and Cohen’s κ 0.237 (95% bootstrap CI 0.174 to 0.304). Brier score was 0.3621 (null Brier 0.0472) and expected calibration error was 0.564, confirming calibration failure due to score compression (range 0.52–0.72) despite preserved discrimination. The sensitivity estimate should be interpreted with caution given the relatively small number of reference-standard positives (n=41); the Wilson CI width of 14.8 percentage points (83.9–98.7%) reflects substantial uncertainty around this point estimate.

Cross-validated results

10-fold cross-validation yielded bias-corrected sensitivity 95.12% (95% Wilson CI 83.9% to 98.7%; optimism 0.00 pp) and specificity 75.80% (95% Wilson CI 72.7% to 78.7%; optimism+1.40 pp), confirming by internal validation that primary metrics are not materially inflated by circular optimisation; independent external validation was not performed.

Conclusions

The DenseNet-121 algorithm demonstrated high point-estimate sensitivity and excellent discrimination for chest radiograph triage in a Nepali health-assessment population, supporting its potential as a radiographic abnormality rule-out triage tool (NPV 99.67%); this does not constitute microbiological exclusion of active pulmonary tuberculosis. Systematic score compression—preserved discrimination despite calibration shift—is a quantifiable marker of low- and middle-income country distributional shift. Prospective local calibration studies and independent external validation are warranted before operational deployment.