Artificial intelligence-supported double reading in European population breast cancer screening: A systematic review and meta-analysis of prospective programs.
Authors
Affiliations (3)
Affiliations (3)
- Division of Breast Imaging, Department of Radiology, University of North Carolina at Chapel Hill, Chapel Hill, NC, USA. Electronic address: [email protected].
- Carolina Mammography Registry, University of North Carolina at Chapel Hill, Chapel Hill, NC, USA.
- Division of Breast Imaging, Department of Radiology, University of North Carolina at Chapel Hill, Chapel Hill, NC, USA.
Abstract
Most European population mammography screening programs rely on double reading with arbitration, a model that delivers mortality benefit but is increasingly challenged by radiologist workload, variable specificity, and interval cancers. Artificial intelligence (AI) is being evaluated to support or optimize these established European screening pathways. To synthesize prospective or program-embedded evaluations of AI conducted within European-style population screening programs and to estimate exploratory program-level absolute risk differences (RDs) per 1000 examinations for cancer detection rate (CDR) and recall. We performed a prespecified, focused evidence synthesis of three large studies embedded within routine population screening programs operating under European-relevant workflows: MASAI (randomized AI-supported risk triage within a national program), ScreenTrustCAD (prospective paired-reader evaluation with AI as an independent reader in a double-reading framework), and PRAIM (nationwide decision-referral implementation). Outcomes were harmonized as AI-control RDs per 1000 examinations. Random-effects pooling used Hartung-Knapp-Sidik-Jonkman models. For the paired-reader design, sensitivity analyses applied a Kish effective sample-size approach across plausible within-examination correlations (ρ = 0.3-0.8). Positive predictive value (PPV) and workflow/time outcomes were summarized descriptively. Across 597,419 examinations, the pooled CDR RD was +0.9 per 1000 (95% CI -0.0 to +1.8; I<sup>2</sup> ≈ 12%), consistent with a modest directional increase with borderline statistical uncertainty. The pooled recall RD was -0.6 per 1000 (95% CI -3.1 to +2.1; I<sup>2</sup> ≈ 41-43%), indicating no consistent recall increase across screening programs. Where reported, PPV was higher with AI-supported screening. Efficiency signals included 44.3% fewer total readings in MASAI and shorter reading times for AI-normal examinations in PRAIM; in PRAIM, a program-level safety-net mechanism recovered 204 cancers that would otherwise have been missed. In European population screening programs characterized by double reading and arbitration, prospective program-embedded evidence suggests that AI integration may yield a small absolute increase in cancer detection (≈1/1000) without a consistent increase in recall, alongside improved PPV and efficiency signals. These findings suggestAI primarily as a complementary reader within European screening workflows, with implementation requiring explicit quality assurance and monitoring of interval cancers and stage distribution.