A Review of Deep Learning Methods for Multimodal Medical Image Fusion.
Authors
Affiliations (1)
Affiliations (1)
- Department of Management Science and Technology, University of Patras, 26334 Patras, Greece.
Abstract
Multimodal medical image fusion (MMIF) aims to integrate complementary information from different imaging modalities into a single, more informative image to support clinical diagnosis. Since 2017, a growing number of deep learning-based approaches have been proposed for MMIF, including various network architectures designed to enhance visual quality. However, a comprehensive and up-to-date review of deep learning-based MMIF techniques is still lacking. To fill this gap, this paper provides a comprehensive survey of deep learning-based MMIF methods. First, we categorize MMIF approaches based on deep learning frameworks and conduct an in-depth analysis of loss functions, evaluation metrics, medical imaging modality pairs, and medical datasets. Then, we review the mainstream deep learning-based MMIF methods, including CNNs, Autoencoders, GANs, and Transformers. Subsequently, we introduce emerging deep learning methods, including diffusion models and Mamba-based methods. In addition, we summarize widely used datasets and evaluation metrics, conduct quantitative experiments on representative methods, and propose a unified set of evaluation metrics for standardized comparison. Finally, we identify key research challenges and outline promising future directions for deep learning-based MMIF. This survey aims to provide researchers with a clear understanding of recent progress in deep learning-based MMIF and to facilitate further studies in this area.