Segmentation-based deep learning emphysema quantification using chest CT: improved accuracy and robustness vs LAA-950.
Authors
Affiliations (9)
Affiliations (9)
- Department of Radiology, Duke University, Durham, NC, USA.
- Department of Electrical and Computer Engineering, Duke University, Durham, NC, USA.
- Department of Medicine, Duke University School of Medicine, Durham, NC, USA.
- Department of Radiology, National Jewish Health, Denver, CO, USA.
- Medical Physics Graduate Program, Duke University, Durham, NC, USA.
- Departments of Physics and Biomedical Engineering, Duke University, Durham, NC, USA.
- Department of Radiology, Duke University, Durham, NC, USA. [email protected].
- Department of Electrical and Computer Engineering, Duke University, Durham, NC, USA. [email protected].
- Medical Physics Graduate Program, Duke University, Durham, NC, USA. [email protected].
Abstract
To develop a deep learning segmentation algorithm to enable accurate, reliable emphysema quantification while improving agreement with radiologist assessments and pulmonary function tests. The model was developed using retrospective virtual and clinical datasets. Virtual data enabled pre-training using ground truth across controlled parameters, including scanners, doses, and reconstruction kernels, while clinical data enabled fine-tuning with expert-annotated emphysema masks. Segmentation accuracy was quantified using the Dice coefficient, and robustness was quantified using emphysema percentage consistency across imaging conditions. Model-based emphysema percentage was correlated with Fleischner visual scores (ordinal; 0-5) and pulmonary function tests (DLCO, FEV<sub>1</sub>pp, and FEV1/FVC) and compared to LAA-950. Statistical analysis included univariate/multivariate correlations. Quantitative assessment included Dice, bias, limits of agreement, and reproducibility coefficient. Virtual data included 20 human models (mean age: 43 years ± 11 [SD], 10 men), and clinical data included multi-center cohorts of 101 patients (C1; 57 years ± 8, 54 men), 23 patients (C2; 57 years ± 7, 14 men), and 1159 patients (C3; 65 years ± 9, 586 men). The model outperformed LAA-950 in segmentation accuracy, achieving higher Dice scores across virtual (76.6% ± 8.8 vs 51.5% ± 23.5), C1 (48.4% ± 24.2 vs 22.5% ± 20.2), and C2 (64.8% ± 9.3 vs 32.7% ± 16.8) cohorts. Furthermore, analysis showed improved bias and limits of agreement (C1; 1.0% ± 4.3 vs 2.1% ± 15.1) and stronger correlations with visual scoring (C1; 0.77 vs 0.47) and pulmonary function tests (C3; DLCO multivariate R²: 0.31-0.32 vs 0.21-0.25). The proposed model provided a promising, clinically relevant alternative to LAA-950, improving accuracy and consistency in emphysema quantification and further aligning with radiologist assessments and lung function metrics. Question Can deep learning-based emphysema segmentation on CT improve the accuracy and robustness of emphysema quantification relative to the traditional LAA-950 biomarker? Findings The deep learning model outperformed LAA-950 in terms of accuracy and reproducibility and demonstrated stronger correlations with visual emphysema scores and pulmonary function tests. Clinical relevance The improved performance of the deep learning approach enables more reliable emphysema quantification across imaging conditions, with closer agreement to radiologist assessments and pulmonary function tests.