Attention-Driven CNNs as a Strong Default for HER2 Prediction from DCE-MRI: A Comparison with Transformer Architectures.
Authors
Affiliations (1)
Affiliations (1)
- Department of Industrial Engineering, Ariel University, Ariel 40700, Israel.
Abstract
<b>Background:</b> HER2 status guides targeted therapy in breast cancer but is currently determined by invasive biopsy. Imaging-based HER2 prediction from dynamic contrast-enhanced MRI (DCE-MRI) could provide a non-invasive adjunct decision-support signal, but published models are typically single-center with heterogeneous preprocessing that limits reproducibility. <b>Methods:</b> We trained a Triple-Head Dual-Attention ResNet (THDA-ResNet) that processes three DCE phases (pre-contrast, early post-contrast, and late post-contrast) on the multicenter BreastDCEDL dataset (<i>n</i> = 1149, I-SPY trials), and we compared it with Vision Transformer (ViT) and Convolutional Vision Transformer (CvT) baselines, all ImageNet-pretrained. We benchmarked 14 preprocessing strategies, with and without N4 bias-field correction. External validation used the independent BreastDCEDL_AMBL cohort (43 lesions). AUC confidence intervals used stratified bootstrap; model comparisons used DeLong's test. <b>Results:</b> THDA-ResNet achieved the highest AUC, 0.74 (95% CI 0.65-0.83), versus 0.66 for ViT and 0.63 for CvT, with the advantage reaching borderline significance over CvT (p=0.054) and not significant over ViT (p=0.14). At a threshold of 0.7, it retained discrimination (sensitivity 0.41, specificity 0.86), while transformers collapsed to near-trivial classifiers. External AUC was 0.66 (0.49-0.81). N4 correction did not improve performance. <b>Conclusions:</b> Attention-driven CNNs are a strong default for HER2 prediction from DCE-MRI on medium-sized cohorts, and N4 correction can be omitted, simplifying the pipeline.