Comparison of Artificial Intelligence Detection and Short-Term Breast Cancer Risk Models within and beyond Their Intended Use in Mammography Screening.
Authors
Affiliations (5)
Affiliations (5)
- Department of Imaging and Pathology, KU Leuven, Leuven, Belgium.
- Department of Radiology, University Hospital Leuven, Herestraat 49, Box 7003, 3000 Leuven, Belgium.
- Faculty of Mathematics and Physics, University of Ljubljana, Ljubljana, Slovenia.
- Jožef Stefan Institute, Ljubljana, Slovenia.
- Department of Medical Physics, University of Wisconsin-Madison, Madison, Wis.
Abstract
Purpose To compare artificial intelligence-based breast cancer detection and short-term breast cancer risk prediction models in their intended and nonintended settings over a single mammography screening round. Materials and Methods In this retrospective study, the Cohort of Screen-Age Women-Case Control dataset from Sweden (May 2008-December 2016; Hologic) was used, including patients with screen-detected and interval cancers, with interval cancers defined as a diagnosis more than 60 days after screening. Two examination-based settings were evaluated: detection (screening examinations) and short-term risk (screen-negative examinations). Mirai Risk, RSNA Detection (2023 challenge winner), Transpara Detection, and Transpara Risk models were assessed in both intended and nonintended settings. Discriminative performance was evaluated using the area under the receiver operating characteristic curve (AUC), and then clinically relevant sensitivity and specificity thresholds were compared. Results A total of 20 187 examinations (7430 women) were included, with 741 examinations leading to diagnosis within 2 years (524 screen-detected and 217 interval cancers) and 19 446 examinations without cancer. Transpara Detection and Transpara Risk achieved similar performance for detection (AUC for each, 0.92; 95% CI: 0.91, 0.93; <i>P</i> = .93) and outperformed the other models (all <i>P</i> < .001). Transpara Detection demonstrated the highest specificity (97.1%; 95% CI: 96.8, 97.3) at double-reading sensitivity (all <i>P</i> < .001). There was no evidence of a difference in specificity of Transpara Risk and RSNA Detection at this sensitivity (<i>P</i> = .29). For risk, Transpara Risk performed best (AUC, 0.81; 95% CI: 0.78, 0.84), with 49.8% sensitivity at 90% specificity (all <i>P</i> ≤ .002) for interval cancers. There was no evidence of a difference in AUC between Mirai Risk and Transpara Detection (<i>P</i> = .96). Conclusion Artificial intelligence-based mammographic detection and short-term breast cancer risk prediction models performed best in their intended settings. <b>Keywords:</b> Mammography, Breast, Computer Applications, Detection/Diagnosis, Screening, Technology Assessment, Model Validation <i>Supplemental material is available for this article.</i> © The Author(s) 2026. Published by the Radiological Society of North America under a CC BY 4.0 license.