Back to all papers

D-TICNN: A dual encoder network with simplified transformer-involution and U-Net-based decoder for enhanced polyp and skin cancer segmentation.

September 17, 2026pubmed logopapers

Authors

Oukdach Y,Garbaz A,Kerkaou Z,El Ansari M,Koutti L,Fouad El Ouafdi A,Salihoun M

Affiliations (7)

  • LabSIV, Department of Computer Science, Faculty of Sciences, Ibnou Zohr University, B.P 8106, Agadir, 80000, Morocco. Electronic address: [email protected].
  • LabSIV, Department of Computer Science, Faculty of Sciences, Ibnou Zohr University, B.P 8106, Agadir, 80000, Morocco. Electronic address: [email protected].
  • LabSIV, Department of Computer Science, Faculty of Sciences, Ibnou Zohr University, B.P 8106, Agadir, 80000, Morocco. Electronic address: [email protected].
  • Informatics and Applications Laboratory, Department of Computer Sciences, Faculty of Science, Moulay Ismail University, B.P 11201, Meknés, 52000, Morocco. Electronic address: [email protected].
  • LabSIV, Department of Computer Science, Faculty of Sciences, Ibnou Zohr University, B.P 8106, Agadir, 80000, Morocco. Electronic address: [email protected].
  • LabSIV, Department of Computer Science, Faculty of Sciences, Ibnou Zohr University, B.P 8106, Agadir, 80000, Morocco. Electronic address: [email protected].
  • Ibn Sina Hospital, Mohammed V University of Rabat, Rabat, 10000, Morocco. Electronic address: [email protected].

Abstract

Medical image segmentation is critical for accurate diagnosis and treatment planning. Deep learning methods, particularly CNNs, have shown promising results in enhancing segmentation accuracy. Transformers, while demonstrating strong potential in disease segmentation, tend to be computationally intensive and complex for clinical applications due to the multi-head attention mechanism in their architecture. In this paper, we propose D-TICNN: a Dual Encoder Architecture with Simplified Transformer-Involution and a U-Net-Based Decoder, designed to improve segmentation performance for polyp and skin cancer images. The dual-encoder structure employs a simplified transformer without attention, reducing computational complexity by focusing on layer normalization and MLP modules, while involution layers enhance local feature learning. Features extracted from the dual encoders are refined through a custom Convolution-Involution Refinement Attention mechanism, which combines the strengths of involution and convolution operations. The U-Net-based decoder incorporates multiple convolutional blocks for feature resampling, supplemented by an enhanced statistical features block that computes the min, max, standard deviation, and mean, which are fused with the decoder output for final mask prediction. We conducted extensive experiments on seven public datasets. For polyp segmentation, CVC-ClinicDB and Kvasir-SEG were used for training, while CVC-ColonDB, Etis-LaribDB, and CVC-300 were used for generalization testing. For skin cancer segmentation, ISIC 2017 was used for training and ISIC 2016 for evaluation. The proposed D-TICNN achieves state-of-the-art results, with a mean Dice coefficient of 93% for polyp segmentation and 92% for skin cancer segmentation.

Topics

Journal Article

Ready to Sharpen Your Edge?

Subscribe to join 11k+ peers who rely on RadAISlice. Get the essential weekly briefing that empowers you to navigate the future of radiology.

We respect your privacy. Unsubscribe at any time.