Multimodal DCE-MRI stacking fusion for luminal versus non-luminal classification in breast cancer: A multicenter study with interpretability and ablation analysis.
Authors
Affiliations (4)
Affiliations (4)
- Department of Radiology, The First Affiliated Hospital of Yangtze University, 8 Hangkong Road, Shashi District, Jingzhou City, Hubei Province, PR China.
- Department of Radiology, Zhuzhou Central Hospital, 116 South Changjiang Road, Taishan Road Subdistrict, Tianyuan District, Zhuzhou City, Hunan Province, PR China.
- Department of Radiology, The First Affiliated Hospital of Yangtze University, 8 Hangkong Road, Shashi District, Jingzhou City, Hubei Province, PR China. Electronic address: [email protected].
- Department of Radiology, The First Affiliated Hospital of Yangtze University, 8 Hangkong Road, Shashi District, Jingzhou City, Hubei Province, PR China. Electronic address: [email protected].
Abstract
To develop and externally validate a multimodal stacking fusion model based on dynamic contrast-enhanced MRI (DCE-MRI) for classifying luminal versus non-luminal breast cancer, and to assess each modality's contribution through interpretability and ablation analyses. We retrospectively enrolled 396 patients with pathologically confirmed invasive breast cancer from two centers, split into training (n = 220), validation (n = 56), and external test (n = 120) sets. Four unimodal models were trained independently on different feature sources: a clinical multilayer perceptron (Clinical-MLP), a whole-tumor radiomics model (Radiomics-ExtraTrees), a subregion-based habitat radiomics model (Habitat-ExtraTrees), and a transfer learning model (DL-ResNet50). Their predicted probabilities were fused via stacking, with the meta-classifier selected through a grid search over nine candidate model families that identified a support vector machine. Modality contributions were quantified by SHapley Additive exPlanations (SHAP) attribution and leave-one-modality-out ablation experiments. The stacking model achieved AUCs of 0.906, 0.850, and 0.854 on the training, validation, and external test sets, It reached the highest AUC, with significant improvements over Clinical-MLP (P < 0.001) and DL-ResNet50 (P = 0.048), and a sensitivity of 0.839 for identifying luminal cases. SHAP ranked DL-ResNet50 as the dominant contributor (48.6%), and ablation confirmed that its removal caused the largest AUC decline (-0.033). Removing Radiomics marginally increased AUC (+0.003). The model achieved robust cross-center performance for luminal versus non-luminal classification. Combining SHAP with ablation revealed that a modality's attribution weight does not guarantee its irreplaceability, suggesting that both analyses are needed to evaluate modality contributions and guide efficient model design.