Diagnostic performance of radiomic features from FS-T2WI for parotid tumor classification: A comparison of deep learning and machine learning models.
Authors
Affiliations (5)
Affiliations (5)
- Department of Oral and Maxillofacial Surgery, The Affiliated Huaian No.1 People's Hospital of Nanjing Medical University, Huaian, Jiangsu Province, 223300, China. Electronic address: [email protected].
- Department of Oral and Maxillofacial Surgery, The Affiliated Huaian No.1 People's Hospital of Nanjing Medical University, Huaian, Jiangsu Province, 223300, China. Electronic address: [email protected].
- Department of Oral and Maxillofacial Surgery, The Affiliated Huaian No.1 People's Hospital of Nanjing Medical University, Huaian, Jiangsu Province, 223300, China. Electronic address: [email protected].
- Department of Oral and Maxillofacial Surgery, The Affiliated Huaian No.1 People's Hospital of Nanjing Medical University, Huaian, Jiangsu Province, 223300, China. Electronic address: [email protected].
- Department of Radiology, The Affiliated Huaian No.1 People's Hospital of Nanjing Medical University, Huaian, Jiangsu Province, 223300, China. Electronic address: [email protected].
Abstract
Accurate discrimination of benign and malignant parotid tumors prior to surgery is of great clinical significance. Conventional magnetic resonance imaging (MRI) still has limitations in terms of objective quantification, and the optimal radiomic modeling strategy for parotid lesions has yet to be established. We extracted radiomic features from fat-suppressed T2-weighted (FS-T2WI) images of 102 patients with pathologically confirmed parotid tumors and constructed four machine learning models (logistic regression, gradient boosting, random forest, and LightGBM) as well as four tabular deep learning models (TabResNet, DCNv2, AutoInt, and TabNet). Model performance was evaluated on an independent held-out test set, with five-fold cross-validation applied only for optimal regularization parameter selection during feature screening. The area under the receiver operating characteristic curve (AUC) as the primary metric and decision curve analysis (DCA) were used to assess clinical utility. On the independent test set, the TabResNet model achieved the highest AUC of 0.8913, which outperformed the best-performing conventional machine learning model (Random Forest, AUC = 0.8641), demonstrating stronger generalization ability and better clinical utility. The other three deep learning architectures provided acceptable discriminatory power but did not yield sustained clinical benefits. Furthermore, a multi-reader trial confirmed that TabResNet model assistance improved diagnostic performance and reading efficiency for clinicians with different seniority and professional backgrounds. FS-T2WI radiomics can reliably and non-invasively distinguish between benign and malignant parotid tumors, while the TabResNet framework offers a practical reference for building high-performance radiomics models in clinical practice.