Back to all papers

Generalizable multi-modal medical image segmentation model using Multi-head Gated Cross Attention fusion Encoder-based adaptive Trans-Mobile-Unet++ with consistency loss function.

August 29, 2026pubmed logopapers

Authors

Neeraja C,Reddy GU,Thumbur G

Affiliations (3)

  • Research Scholar, Department of ECE, Sri Venkateswara University College of Engineering, Sri Venkateswara University, Tirupati, Andhra Pradesh 517502, India. Electronic address: [email protected].
  • Professor, Department of Electronics and Communication Engineering, Sri Venkateswara University College of Engineering, Sri Venkateswara University, Tirupati, Andhra Pradesh, 517502, India. Electronic address: [email protected].
  • Associate Professor, Department of EECE, GITAM School of Technology, GITAM(Deemed to be University), Visakhapatnam, Andhra Pradesh, 530045, India. Electronic address: [email protected].

Abstract

Automated medical image segmentation is an essential component of modern clinical diagnosis, enabling the precise delineation of anatomical features and abnormal tissue from various medical imaging techniques, including Magnetic Resonance Imaging (MRI) and Computed Tomography (CT) scans. However, achieving reliable segmentation across multi-modal medical images remains difficult because of variations in imaging properties, feature distributions, and modality-specific information. Existing segmentation methods often exhibit limited generalization and insufficient cross-modal feature integration, resulting in reducing segmentation accuracy. To resolve these concerns, the proposed framework employs an advanced attention mechanism-driven deep learning network to effectively learn complex heterogeneous features in medical images. The process begins by collecting the required multi-modal images from standard public databases, which are then directly provided to the segmentation model for performing multi-modal segmentation. The segmentation process is executed using Multi-head Gated Cross Attention Fusion Encoder-based Adaptive Transformer Mobile-UNet++ with Consistency Loss Function (MGAFE-ATMU++-CLF) for accurate segmentation of multi-modal medical images. The proposed model incorporates a multi-head gated cross-attention fusion mechanism throughout the encoder to efficiently extract complementary information from different imaging modalities while preserving discriminative features. Furthermore, a Renovated Puma Optimizer (RPO) is designed to fine-tune the critical network hyperparameters for the MGAFE-ATMU++ framework, thereby enhancing the segmentation capability of the developed approach. The evaluation results indicate the introduced MGAFE-ATMU++-CLF framework delivers improved segmentation accuracy by effectively extracting heterogeneous features, integrating cross-modal information, and accurately delineating anatomical boundaries, highlighting its potential for reliable analysis of multi-modal medical images.

Topics

Journal Article

Ready to Sharpen Your Edge?

Subscribe to join 11k+ peers who rely on RadAISlice. Get the essential weekly briefing that empowers you to navigate the future of radiology.

We respect your privacy. Unsubscribe at any time.