Performance of Machine Learning Models Based on Medical Imaging in Predicting Pathological Grade of Clear Cell Renal Cell Carcinoma.
Authors
Affiliations (7)
Affiliations (7)
- Department of Urology, Second Affiliated Hospital of Dalian Medical University, Dalian, Liaoning, China.
- Liaoning Provincial Key Laboratory of Urological Digital Precision Diagnosis and Treatment, Dalian, Liaoning, China.
- Liaoning Engineering Research Center of Integrated Precision Diagnosis and Treatment Technology for Urological Cancer, Dalian, Liaoning, China.
- Dalian Key Laboratory of Prostate Cancer Research, Dalian, Liaoning, China.
- Second Clinical College, Dalian Medical University, Dalian, Liaoning, China.
- Department of Urology, Beijing United Family Hospital and Clinics, Beijing, China.
- Center for Evidence-Based and Translational Medicine, Hubei Key Laboratory of Urinary System Diseases, Zhongnan Hospital of Wuhan University, Wuhan, China.
Abstract
Predicting clear cell renal cell carcinoma (ccRCC) pathological grade preoperatively is critical for clinical management. This study aims to evaluate the diagnostic accuracy and clinical utility of machine learning (ML)-based imaging models. The Cochrane Library, PubMed, Embase, and Scopus databases were searched systematically for studies published before January 2026. Study quality was assessed using the Quality Assessment of Diagnostic Accuracy Studies 2 and the Radiomics Quality Score. Pooled sensitivity, specificity, positive/negative likelihood ratios (PLRs/NLRs), diagnostic scores, diagnostic odds ratios (DORs), and summary receiver operating characteristic (SROC) curves were calculated. Decision curve analysis (DCA) and Fagan nomogram analysis were performed to evaluate clinical utility. Subgroup analyses were conducted to further explore sources of heterogeneity. A total of 43 studies involving 12,675 patients were included. The area under the SROC curve was 0.89, with a sensitivity of 0.79, specificity of 0.85, PLR of 5.27, NLR of 0.25, diagnostic score of 3.07, and DOR of 21.52. Fagan analysis revealed a positive prediction increased the posttest probability of high-grade disease to 70%, whereas a negative prediction decreased it to 10%. DCA demonstrated a net benefit over standard strategies across a 0.10-0.70 threshold range. Subgroup analyses revealed significantly greater sensitivity for the deep learning (DL) models than for the radiomics (0.91 vs. 0.75; p < 0.01) and automatic models compared with manual segmentation (0.86 vs. 0.76; p = 0.03). Notably, single-center independent validation (0.92) outperformed both multicenter external (0.79) and internal validation (0.72) strategies (p < 0.01). No significant performance differences were observed across imaging modalities, phase protocols, clinical variable integration, geographic regions, or sample sizes. This study confirms the significant potential of radiomics and DL models for the preoperative prediction of the pathological grade of ccRCC. Nevertheless, future multicenter validation is essential to address the performance gap observed in external datasets. Prospero: CRD42023455847.