Back to all papers

Augmenting medical image specific foundation model with classical radiomic signatures for improved nasopharyngeal carcinoma stage classification.

July 15, 2026pubmed logopapers

Authors

Li Y,Liu J,Wen X,Niu X,Wei L,Yuan J,Wang Z,Na R

Affiliations (5)

  • Department of Otolaryngology-Head and Neck Surgery, The Eighth Affiliated Hospital, Sun Yat-Sen University, Shenzhen, China.
  • Department of Urology, Ruijin Hospital, Shanghai Jiao Tong University School of Medicine, Shanghai, China.
  • Department of Orthopedics, Shanghai Fengxian District Central Hospital, Shanghai Jiao Tong University Affiliated Sixth People's Hospital South Campus, Shanghai, China.
  • Department of Surgery, School of Clinical Medicine, LKS Faculty of Medicine, The University of Hong Kong, Hong Kong, Hong Kong SAR, China.
  • Department of Surgery, Queen Mary Hospital, Hong Kong, Hong Kong SAR, China.

Abstract

Accurate clinical staging is the critical process of prognosis and treatment planning for nasopharyngeal carcinoma (NPC). Conventional radiomics is limited to hand-crafted features that may overlook complex patterns, while generic deep learning models often lack domain specificity. Currently, SAM-Med3D, a vision transformer (ViT)-based medical model excels at capturing global volumetric context. We hypothesized that fusing radiomic and SAM-Med3D feature-extraction would yield synergistic value. This study aims to construct an MRI-based fusion model for NPC stage classification (Stage I, II and III vs. Stage IVa) and benchmark its performance against experienced experts. We retrospectively analyzed 264 patients with non-metastatic nasopharyngeal carcinoma, partitioned into training (<i>n</i> = 198) and independent validation (<i>n</i> = 66) cohorts. MRI sequences (T1WI, T2WI, CE-T1WI) were utilized to extract high-dimensional features by two distinct paradigms: hand-crafted radiomics and deep features from SAM-Med3D. An early-fusion dataset (SAM-PR) integrating both feature types was also constructed. Following LASSO-based dimensionality reduction, five machine learning classifiers (LR, SVM, RF, XGB, LGBM) were developed for staging. Model performance was evaluated using the area under the curve (AUC) and benchmarked against the independent, blinded staging of two senior otolaryngologists. Optimism-Corrected Bootstrap analysis was applied to evaluate the reliability of AUC, and the learning curve was conducted to avoid overfitting. LASSO regression identified an optimal fusion set of 17 features (9 hand-crafted and 8 deep-learned). Among all 15 constructed DL models, the fusion-based LR yielded the optimal performance (AUC = 0.760, optimism-corrected AUC of 0.836), and the fusion approaches consistently yielded higher AUCs than individual feature approaches. Compared to human experts, our optimal model achieved a more balanced performance (Sensitivity: 73.91% vs. 88%; Specificity: 69.77% vs. 26%-30%) and a higher F1-score (0.615 vs. 0.353-0.400). The integrated model, combining conventional radiomics with SAM-Med3D features, consistently showed higher accuracy than its unimodal counterparts across five distinct algorithms (LR, SVM, RF, XGB, LGBM) for NPC stage I, II, III and stage IVa differentiation. All five fusion models demonstrated comparable discriminative ability compared to experienced experts. These findings suggest that hybrid feature frameworks can serve as powerful objective tools in NPC staging.

Topics

Journal Article

Ready to Sharpen Your Edge?

Subscribe to join 11k+ peers who rely on RadAI Slice. Get the essential weekly briefing that empowers you to navigate the future of radiology.

We respect your privacy. Unsubscribe at any time.