Data-centric strategies enable cross-center generalization of convolutional neural networks for CT brain extraction.
Authors
Affiliations (3)
Affiliations (3)
- School of Health Science and Engineering, University of Shanghai for Science and Technology, Shanghai, China.
- Institute of Diagnostic and Interventional Radiology, Shanghai Sixth People's Hospital Affiliated to Shanghai Jiao Tong University School of Medicine, Shanghai, China.
- Department of Radiology, Liyang Hospital of Chinese Medicine, Changzhou, Jiangsu, China.
Abstract
To determine if a small number of annotated scans can enable clinically meaningful adaptation of a pre-trained convolutional neural network across different centers. Using non-contrast head CT brain extraction as the test case, we also sought to identify the factors contributing to performance drops when applying models to new sites. Non-contrast head CT scans were used in this study, from two separate institutions (Center A: 595 scans; Center B: 486 scans). A public pre-trained U-Net model for brain tissue extraction was first evaluated on Center A dataset as a baseline. The model was then fine-tuned incrementally using stratified subsets of annotated scans (5, 15, 25, and 75). The pre-trained model was subsequently transferred to Center B dataset and adapted through the same incremental fine-tuning strategy. The primary metric of interest was the three-dimensional Dice similarity coefficient (DSC), with a pre-defined minimal important difference (MID) set at 0.01. To identify predictors of baseline DSC and fine-tuning gain driving cross-center performance degradation, variance inflation factor (VIF) screening and logit Gaussian linear models (GLMs) were applied, with SHapley Additive exPlanations (SHAP) analysis providing qualitative feature importance rankings. At center A, the pre-trained model achieved a DSC of 0.9634 ± 0.0151. Fine-tuning with five annotated scans raised DSC to 0.9784 ± 0.0070 (<i>p</i> < 0.0001), exceeding the MID by 1.5-fold and reducing clinically unacceptable segmentations from 8.7 to 0.8%. A comparable gain was observed at Center B (<i>p</i> < 0.0001). Further increases to 15, 25, and 75 samples offered only marginal improvement below the MID threshold. Regression and SHAP analyses identified slice thickness as the dominant driver of both baseline performance degradation and fine-tuning improvement, accounting for 80-90% of predictive weight at both centers. For NCCT brain extraction assessed across two centers, fine-tuning a publicly pre-trained CNN using five local annotated scans can generate clinically valuable cross-center segmentation gains; slice thickness heterogeneity appears to be the dominant source of inter-center domain shift.