Local adaptation of a chest radiograph AI model for pulmonary tuberculosis triage in Uzbekistan: Pooled and site-stratified internal validation and exploratory workstation integration.
Authors
Affiliations (4)
Affiliations (4)
- New Uzbekistan University, Tashkent, Uzbekistan. Electronic address: [email protected].
- Alfraganus University, Tashkent, Uzbekistan.
- New Uzbekistan University, Tashkent, Uzbekistan.
- Jizzakh State Pedagogical University, Jizzakh, Uzbekistan.
Abstract
Dataset shift can limit the transferability of chest radiograph artificial-intelligence models. To assess pooled and site-stratified internal performance after local adaptation of a public pulmonary tuberculosis (TB) classifier in Uzbekistan. The cohort included 502 unique patients: 96 with specialist-assigned clinical-radiological TB-positive status and 406 with TB-negative status. The radiologist and phthisiatrician initially assessed each case independently using chest radiography, CT, and available contemporaneous clinical information, then established the final reference classification by consensus without access to the AI output. The primary evaluation used a fixed stratified 70 %/15 %/15 % train/validation/held-out split; a complementary five-fold out-of-fold (OOF) analysis generated predictions across all 502 patients. Complete standardized microbiological confirmation was unavailable. On the primary held-out subset (n = 76; 15 TB-positive), the adapted model achieved ROC AUC 0.9902 (95 % CI 0.9694-1.0000), F1 0.9375 (0.8571-1.0000), and precision-recall AUC 0.9557 (0.8717-1.0000). The paired AUC difference from the public baseline was 0.0054 (95 % CI -0.0077 to 0.0230; p = 0.4864), showing no statistical superiority. Given only 15 TB-positive held-out cases, these high point estimates remain preliminary and imprecise. Adapted held-out AUCs were 0.9844 (0.9375-1.0000) at the District Hospital and 0.9973 (0.9841-1.0000) at the TB Dispensary; these were site-stratified internal estimates, not institution-held-out validation. Complementary aggregated OOF AUC was 0.9772. The locally adapted model showed strong internal discrimination and was integrated into a research prototype, but statistical superiority over the public baseline was not established. The diagnostically enriched cohort, small held-out set, specialist clinical-radiological reference, absence of institution-held-out validation, and uncontrolled workflow observations preclude claims of population-screening performance, microbiologically confirmed diagnostic accuracy, external generalizability, or improved clinical efficiency.