Back to all papers

Parotid gland tumor computed tomography image segmentation and benign-malignant differentiation diagnosis.

August 5, 2026pubmed logopapers

Authors

Song N,Sun C,Li Q,He S,Song J,Yan K,Zhang Y,Wang W,Cai K,Fu T,Qiu Z,Piao S

Affiliations (5)

  • Department of Oral and Maxillofacial Surgery, The First Affiliated Hospital of Harbin Medical University, Harbin, China.
  • Department of Oral and Maxillofacial Surgery, Qunli Branch Campus, The First Affiliated Hospital of Harbin Medical University, Harbin, China.
  • Outpatient CT Room, The First Affiliated Hospital of Harbin Medical University, Harbin, China.
  • School of Stomatology, Harbin Medical University, Harbin, China.
  • College of Computer and Control Engineering, Northeast Forestry University, Harbin, China.

Abstract

Parotid gland tumors are the most common type of salivary gland neoplasms. The accurate determination of tumor nature is essential for clinical treatment decision-making. Existing deep learning (DL)-based methods for image segmentation and differential diagnosis often yield blurred boundaries when segmenting parotid tumors on computed tomography (CT) images and face difficulties in distinguishing benign from malignant lesions. This study aimed to develop a solution based on an improved TransUNet to achieve precise tumor segmentation and to enable automatic benign-malignant differentiation of parotid gland tumors by fusing CT images with clinical text information. First, we propose an enhanced PT-TransUNet segmentation model. Built upon the classic TransUNet architecture, it incorporates a learnable Difference of Gaussians (DoG) edge enhancement module at the front end to adaptively sharpen tumor boundary features. In addition, a PSA-CBAM module-which integrates multi-scale convolution with a convolutional block attention module-is embedded in the decoder to improve the model's capability to capture multi-scale features. Second, we construct a multimodal diagnostic model that uses PT-TransUNet as the image branch to extract CT image features, employs the Chinese BERT Whole Word Masking (WWM) model as the text branch to extract clinical text features, and performs deep feature fusion via a bidirectional cross-modal attention mechanism to achieve binary classification of benign versus malignant tumors. On a dataset of 158 CT images, PT-TransUNet achieved Dice coefficients of 0.8212±0.0269 for parotid gland segmentation and 0.8075±0.0245 for parotid tumor segmentation, compared with 0.8118±0.0324 and 0.7928±0.0350 by nnUNet, respectively (P<0.05). The 95% Hausdorff distance (HD95) was 16.14±0.58 mm for benign tumors and 6.28±0.58 mm for malignant tumors. In the multimodal diagnostic task, the fusion model using bidirectional cross-modal attention achieved an accuracy of 0.8909±0.0249 and a sensitivity of 0.9000±0.1369, whereas the CT-only model achieved 0.6091±0.0249 accuracy and 0.1500±0.1369 sensitivity, and the text-only model achieved 0.7273±0.0321 accuracy and 0.4000±0.1369 sensitivity (P<0.001 for both comparisons). Compared with feature-level concatenation, which achieved an accuracy of 0.8455±0.0518, and decision-level weighted average, which achieved an accuracy of 0.7545±0.0249, the cross-modal attention mechanism yielded higher accuracy. This study confirms the superiority of the proposed PT-TransUNet model for CT image segmentation of parotid tumors and highlights the great potential of the cross-modal fusion strategy that integrates CT images with clinical text information for benign-malignant differentiation. This approach provides a new technical pathway for automated and accurate assisted diagnosis of parotid gland tumors and holds important value for clinical application.

Topics

Journal Article

Ready to Sharpen Your Edge?

Subscribe to join 11k+ peers who rely on RadAISlice. Get the essential weekly briefing that empowers you to navigate the future of radiology.

We respect your privacy. Unsubscribe at any time.