Back to all papers

3D ViT network with selective feature enhancement and cross-dimensional interaction for colorectal T-stage classification in CT images.

September 22, 2026pubmed logopapers

Authors

Wang J,Gong X,Jintang R,Chen S,Zheng S,Zhang J

Affiliations (3)

  • College of Physics and Information Engineering, Fuzhou University, Fuzhou, Fuzhou University, No.2 Wulongjiang North Avenue, Fuzhou University Town, Minhou County, Fuzhou City, Fujian Province, China, 350108, Fuzhou, Fujian, 350108, China.
  • Fuzhou University, No.2 Wulongjiang North Avenue, Fuzhou University Town, Minhou County, Fuzhou City, Fujian Province, China, 350108, Fuzhou, Fujian, 350108, China.
  • Fujian Medical University Union Hospital, No. 29, Xinquan Road, Gulou District, Fuzhou City, Fujian Province., Fuzhou, Fujian, 350001, China.

Abstract

Obstructive colorectal cancer (OCC) often manifests on CT as highly variable in morphology and location, with indistinct margins, atypical infiltration signs, and frequent adherence to surrounding structures. In addition, motion artifacts caused by intestinal peristalsis further degrade image quality and interfere with the identification of subtle structures. These factors readily lead to misassessment in T-stage evaluation, and accurately distinguishing between T4 and non-T4 stages is critical as it directly influences surgical planning and clinical decision-making. To address challenges of imaging heterogeneity and motion artifacts in T-stage evaluation, we proposed a 3D vision transformer (ViT) based model with feature enhancement and cross-dimensional interaction for T4/non-T4 classification of OCC. First, we extended ViT into 3D and incorporate a selective feature enhancement module. This module computes attention weights based on feature means, adaptively enhancing discriminative information while suppressing redundant responses. It effectively mitigates difficulties in identifying subtle structures like serosal invasion caused by motion artifacts and low contrast in CT images, enabling the model to focus on diagnostically critical regions. Second, we introduced a cross-dimensional adaptive feature aggregation mechanism that established multi-scale 3D spatial dependencies, integrating global context with fine-grained local details. This significantly improves the model's ability to characterize tumors with complex and variable morphological features, thereby reducing staging inaccuracies due to underutilized spatial information. Furthermore, cross-domain knowledge transfer was employed to enrich the model's representational capacity and alleviate the inherent data scarcity in 3D medical image analysis. Evaluated on a private dataset of 127 CT scans, the proposed model achieved an AUC of 0.8917 and an accuracy of 0.7917 in T4/non-T4 classification, significantly outperforming mainstream 3D medical image classification networks. Our approach offers clinicians a reliable decision support tool for accurate OCC staging, allowing surgeons to tailor operative strategies based on quantitatively reliable imaging assessments.

Topics

Journal Article

Ready to Sharpen Your Edge?

Subscribe to join 11k+ peers who rely on RadAISlice. Get the essential weekly briefing that empowers you to navigate the future of radiology.

We respect your privacy. Unsubscribe at any time.