Deep learning for multimodal brain tumor segmentation: Architectures, fusion, robust learning, and deployment perspectives.
Authors
Affiliations (4)
Affiliations (4)
- Department of Electrical Engineering, Faculty of Engineering, Universiti Malaya, Lembah Pantai, Kuala Lumpur, 50603, Malaysia; Faculty of Public Health, Hubei University of Medicine, 16 Shanghai Road, Shiyan, 442000, Hubei, China. Electronic address: [email protected].
- Department of Electrical Engineering, Faculty of Engineering, Universiti Malaya, Lembah Pantai, Kuala Lumpur, 50603, Malaysia. Electronic address: [email protected].
- Department of Electrical Engineering, Faculty of Engineering, Universiti Malaya, Lembah Pantai, Kuala Lumpur, 50603, Malaysia. Electronic address: [email protected].
- Faculty of Public Health, Hubei University of Medicine, 16 Shanghai Road, Shiyan, 442000, Hubei, China. Electronic address: [email protected].
Abstract
Accurate segmentation of brain tumors from multimodal magnetic resonance imaging (MRI) is essential for diagnosis, treatment planning, radiotherapy targeting, and longitudinal assessment. Deep learning has advanced this task through convolutional neural networks, Transformers, state-space models, diffusion methods, and foundation models. However, strong benchmark performance does not guarantee clinical reliability because systems remain vulnerable to missing or degraded modalities, cross-center shift, and subregion-specific failures, particularly in enhancing tumor (ET) and tumor core (TC). This survey critically reviews deep learning for multimodal brain tumor segmentation from a deployment-oriented, failure-focused perspective. We introduce a unifying framework linking input reliability, fusion-architecture co-design, subregion-specific failure mechanisms, and uncertainty-aware clinical triage. Using this framework, we compare major architectural paradigms in contextual modeling, boundary preservation, computational feasibility, and robustness across ET, TC, and whole tumor (WT). We reinterpret multimodal fusion as a reliability-allocation problem and examine early, intermediate, token-level, sequence-aware, and adaptive strategies under incomplete or degraded inputs. We also synthesize robust learning approaches for sparse supervision, MRI quality degradation, cross-center variation, and test-time adaptation, and assess interpretability, uncertainty estimation, and human-in-the-loop review as mechanisms for clinical risk control. Finally, our evidence-oriented benchmarking analysis identifies WT performance saturation, persistent ET/TC instability, inconsistent boundary-metric reporting, and insufficient stress testing, center-stratified evaluation, calibration assessment, and computational transparency. We conclude that progress should be judged not only by benchmark accuracy but also by subregion-level reliability under realistic deployment conditions.