Back to all papers

Diagnostic accuracy of deep learning-assisted low-dose CT screening for lung cancer detection: a systematic review and meta-analysis.

August 10, 2026pubmed logopapers

Authors

Gebremichael D,Zimme ZA,Mengesha LY,Jibat N,Berihun A,Ayele AA,Ashenafi SA,Tegegn HG,Gete KY

Affiliations (8)

  • School of Public Health, Washington University in St. Louis, St. Louis, MO, USA.
  • College of Health Sciences, Addis Ababa University, Addis Ababa, Ethiopia.
  • Johns Hopkins Bloomberg School of Public Health, Baltimore, MD, USA.
  • School of Health, Faculty of Medicine and Health, University of New England, Armidale, 2351, Australia.
  • Myungsung Medical College, Addis Ababa, Ethiopia.
  • Department of Clinical Pharmacy, School of Pharmacy, College of Medicine and Health Sciences, University of Gondar, Gondar, Ethiopia.
  • School of Medicine, College of Medicine and Health Sciences, Bahir Dar University, Bahir Dar, Ethiopia. [email protected].
  • EPIC Health Systems, Addis Ababa, Ethiopia. [email protected].

Abstract

Deep learning (DL)-assisted low-dose computed tomography (LDCT) may improve lung cancer screening, but the available evidence is heterogeneous and patient-level diagnostic accuracy remains uncertain. MEDLINE, Embase, and Web of Science were searched from January 2010 to December 2025 for studies evaluating DL-assisted LDCT in lung cancer screening or screening-relevant populations. Two reviewers independently screened studies, extracted data, and assessed risk of bias using QUADAS-2. AI-specific reporting completeness was assessed descriptively using items adapted from CLAIM and STARD-AI. Studies with complete or reconstructible patient-level 2 × 2 data at a defined threshold were pooled using bivariate random-effects and hierarchical summary receiver operating characteristic models. Studies without sufficient threshold-specific data were synthesised narratively. Eleven studies met the inclusion criteria. Five studies, comprising 2,220 participants, 232 lung cancer cases, and 1,988 non-cases, provided usable threshold-specific 2 × 2 data and were included in the meta-analysis. Six studies were synthesised narratively because patient-level TP, FP, FN, and TN values were unavailable or not reconstructible. Pooled sensitivity was 85.4% (95% CI, 79.1-90.0), and pooled specificity was 83.5% (95% CI, 73.5-90.2). The positive likelihood ratio was 5.17, the negative likelihood ratio was 0.18, and the diagnostic odds ratio was 29.53. No study was at low risk of bias across all QUADAS-2 domains, and AI-specific reporting gaps were common. Subgroup, meta-regression, and sensitivity analyses were not feasible because only five studies were quantitatively eligible. No statistical evidence of small-study effects was detected, although this assessment was inconclusive because only five studies were pooled. DL-assisted LDCT shows promising but preliminary diagnostic accuracy for lung cancer screening. However, the small, heterogeneous, and methodologically limited evidence base does not support autonomous clinical use. DL is currently best considered a decision-support tool within radiologist-led screening pathways, pending prospective external validation and workflow-based evaluation.

Topics

Journal Article

Ready to Sharpen Your Edge?

Subscribe to join 11k+ peers who rely on RadAI Slice. Get the essential weekly briefing that empowers you to navigate the future of radiology.

We respect your privacy. Unsubscribe at any time.