Deep learning models for pancreatic cancer detection on CT: a meta-analysis.
Authors
Affiliations (1)
Affiliations (1)
- Bishan Hospital of Chongqing Medical University, Chongqing, China.
Abstract
This systematic review and meta-analysis evaluates the diagnostic accuracy of deep learning models for detecting pancreatic ductal adenocarcinoma using computed tomography scans. A systematic literature search of PubMed, Web of Science, and the Cochrane Library was conducted for studies published up to September 11, 2025. Study screening and data extraction were performed independently by two reviewers. Studies utilizing deep learning models for pancreatic ductal adenocarcinoma detection on computed tomography, with histopathological confirmation as the reference standard, were included. Studies employing solely traditional machine learning models were excluded. Primary outcomes were sensitivity, specificity, diagnostic odds ratio, and area under the curve. Quality assessment was performed using standardized tools for risk of bias and methodological soundness. A subgroup analysis focused on scans acquired 3 to 36 months prior to clinical diagnosis. Fifteen studies were included. The pooled results demonstrated a sensitivity of 0.92 (95% confidence interval: 0.88-0.94), specificity of 0.96 (95% CI: 0.92-0.98), diagnostic odds ratio of 285.00 (95% CI, 97.00-839.00), and area under the curve of 0.97 (95% CI, 0.95-0.98). In the preclinical diagnosis subgroup, the pooled sensitivity was 0.73 and specificity was 0.92. This meta-analysis validates the high diagnostic accuracy of deep learning models for detecting pancreatic cancer on CT scans in retrospective datasets. However, all included studies were retrospective, external validation was inconsistent, and no prospective or real-world deployment evidence exists to date. Therefore, while these findings suggest the potential of DL as a decision-support tool, they should be interpreted with caution, and prospective multicenter validation studies are urgently needed before clinical integration can be recommended. However, the pooled estimates should be interpreted cautiously because only one representative model per study was included.