Enhancing Robustness of Deep Learning to Batch Effects from Multi-site Data for Segmentation of Clinically Significant Prostate Cancer on MRI.
Authors
Affiliations (10)
Affiliations (10)
- Department of Biomedical Engineering and Informatics, Luddy School of Informatics, Computing and Engineering, Indiana University Indianapolis, Indianapolis, IN, USA.
- Department of Chemistry, School of Arts and Sciences, University of Pennsylvania, Philadelphia, PA, USA.
- School of Computer Science, Georgia Institute of Technology, Atlanta, GA, USA.
- Wallace H. Coulter Department of Biomedical Engineering, Emory University, Atlanta, GA, USA.
- Department of Urology, Emory University School of Medicine, Atlanta, GA, USA.
- Department of Radiology and Imaging Sciences, Indiana University School of Medicine, Indianapolis, IN, USA.
- Department of Radiology and Imaging Sciences, Emory University, Atlanta, GA, USA.
- Research Career Scientist, Atlanta Veterans Administration Medical Center, Atlanta, USA.
- Department of Urology, Indiana University School of Medicine, Indianapolis, IN, USA.
- Department of Biomedical Engineering and Informatics, Luddy School of Informatics, Computing and Engineering, Indiana University Indianapolis, Indianapolis, IN, USA. [email protected].
Abstract
Deep learning (DL) has shown promise in segmenting clinically significant prostate cancer (csPCa) on MRI. However, batch effects arising from multi-site variations impact their generalizability. This study investigates the impact of training data harmonization, diversity, and few-shot fine-tuning on DL in segmenting csPCa on multi-site MRI. 3 T prostate MRI from N = 1822 patients of four sites (public: D₁, N = 1500; D₂, N = 157; institutional: D₃, N = 47; D₄, N = 118) was leveraged. Bi-parametric MRI (T2-weighted, ADC) was harmonized with csPCa lesions delineated by expert radiologists. nnU-Net DL models for csPCa segmentation on MRI were trained separately (C₁ on D₁; C₂ on D₂) and jointly (C₃ on D₁-D₂). The jointly trained C₃ model was then fine-tuned on D₃ and D₄ using few-shot learning in increments of 50%, 75%, and 100% of train data. Models were primarily evaluated on holdout test sets using sensitivity primarily, in addition to AUC, DSC, precision, and Hausdorff distance (paired t-tests for evaluating significance). Batch effects persisted across datasets despite harmonization. For csPCa segmentation on D<sub>1</sub>/D<sub>2</sub> test sets, C<sub>1</sub> achieved sensitivities of 0.51 ± 0.34/0.11 ± 0.17, C<sub>2</sub> achieved 0.23 ± 0.31/0.17 ± 0.22, while C<sub>3</sub> improved to 0.48 ± 0.36/0.33 ± 0.26. On D<sub>3</sub>/D<sub>4</sub> test sets, C<sub>3</sub> achieved zero-shot sensitivities of 0.19 ± 0.22/0.37 ± 0.35. Few-shot fine-tuning on D<sub>3</sub> and D<sub>4</sub> improved sensitivity to 0.32 ± 0.26/0.54 ± 0.30 on D<sub>3</sub>/D<sub>4</sub> test sets after using 100% of available fine-tuning samples. On subset analyses, lesions > 0.5 cm<sup>3</sup> had consistently higher segmentation performance compared to smaller lesions < 0.5 cm<sup>3</sup>. Diverse multi-site training improves robustness of DL models for csPCa segmentation. Pre-trained models perform modestly on institutional MRI datasets under zero-shot inference, but few-shot fine-tuning on target sites enhances performance.