Diagnostic accuracy of a DenseNet-121 deep learning algorithm for chest radiograph triage in health assessment applicants: a prospective shadow-mode validation study in Nepal.
Authors
Affiliations (2)
Affiliations (2)
- Department of Radiology and Imaging, Patan Academy of Health Sciences, Lalitpur, Bagmati, Nepal [email protected].
- Department of Radiology and Imaging, Patan Academy of Health Sciences, Lalitpur, Bagmati, Nepal.
Abstract
To evaluate the diagnostic accuracy of a publicly available DenseNet-121 convolutional neural network (TorchXRayVision) for triaging chest radiographs of health assessment applicants at a tertiary hospital in Nepal. Prospective, single-centre, shadow-mode diagnostic accuracy validation study. Reported in accordance with the Standards for Reporting of Diagnostic Accuracy Studies (STARD) 2015 checklist and STARD-Artificial Intelligence (AI)/Developmental and Exploratory Clinical Investigations of DEcision support systems driven by Artificial Intelligence (DECIDE-AI) guidelines. Department of Radiology and Imaging, Patan Academy of Health Sciences/Patan Hospital, Lalitpur, Nepal. 826 consecutive health assessment applicants (foreign employment predeparture medical examination and student migration) undergoing chest radiography from 5 June 2026 to 20 June 2026 inclusive (16 days). Two cases were excluded due to Digital Imaging and Communications in Medicine technical failure. DenseNet-121 algorithm (TorchXRayVision library, densenet121-res224-all pretrained weights). A maximum aggregated pathology probability score was derived per radiograph and compared against a post hoc derived threshold of 0.6258 (selected as the highest threshold achieving the prespecified ≥95% sensitivity criterion). Single-reader-per-case review by one of three radiologists-two board-certified radiodiagnosticians (LS: 276 cases; DM: 275 cases) and one radiology resident (UB: 275 cases)-each blinded to AI output, using a standardised data collection worksheet capturing binary classification (abnormal/normal) and free-text findings. Of 826 radiographs, 41 (4.97%) were classified as abnormal by the reference standard. At the post hoc derived threshold of 0.6258, the DenseNet-121 algorithm achieved: sensitivity 95.12% (95% CI 83.9% to 98.7%), specificity 77.2% (95% CI 74.1% to 80.0%), area under the receiver operating characteristic curve 0.9583 (95% bootstrap CI 0.9225 to 0.9843), negative predictive value (NPV) 99.67% (95% Wilson CI 98.8% to 99.9%), positive predictive value 17.89% (95% Wilson CI 13.4% to 23.5%) and Cohen's κ 0.237 (95% bootstrap CI 0.174 to 0.304). Brier score was 0.3621 (null Brier 0.0472) and expected calibration error was 0.564, confirming calibration failure due to score compression (range 0.52-0.72) despite preserved discrimination. The sensitivity estimate should be interpreted with caution given the relatively small number of reference-standard positives (n=41); the Wilson CI width of 14.8 percentage points (83.9-98.7%) reflects substantial uncertainty around this point estimate. 10-fold cross-validation yielded bias-corrected sensitivity 95.12% (95% Wilson CI 83.9% to 98.7%; optimism 0.00 pp) and specificity 75.80% (95% Wilson CI 72.7% to 78.7%; optimism+1.40 pp), confirming by internal validation that primary metrics are not materially inflated by circular optimisation; independent external validation was not performed. The DenseNet-121 algorithm demonstrated high point-estimate sensitivity and excellent discrimination for chest radiograph triage in a Nepali health-assessment population, supporting its potential as a radiographic abnormality rule-out triage tool (NPV 99.67%); this does not constitute microbiological exclusion of active pulmonary tuberculosis. Systematic score compression-preserved discrimination despite calibration shift-is a quantifiable marker of low- and middle-income country distributional shift. Prospective local calibration studies and independent external validation are warranted before operational deployment.