Back to all papers

AI performance varies considerably across mammography devices: a multi-site and multi-vendor retrospective study.

August 26, 2026pubmed logopapers

Authors

Blum M,Morant R,Eichenberger A,Geissler A,Subelack J,Vogel J,Ehlig D

Affiliations (4)

  • Chair of Health Economics, Policy and Management, School of Medicine, University of St.Gallen, St.Gallen, Switzerland. [email protected].
  • Cancer League of Eastern Switzerland, St.Gallen, Switzerland. [email protected].
  • Cancer League of Eastern Switzerland, St.Gallen, Switzerland.
  • Chair of Health Economics, Policy and Management, School of Medicine, University of St.Gallen, St.Gallen, Switzerland.

Abstract

To evaluate an artificial intelligence (AI) algorithm on three different mammography devices and assess screening performance using general and device-specific thresholds. This retrospective study analyzed 22,673 AI-assessed mammographies of women aged 50-69 collected in 2022-2023 from three mammography devices across six screening sites within a Swiss Mammography Screening Program. 115 screen-detected and 15 1-year interval breast cancers were included in the analyses. A commercially available AI algorithm assessed each mammography with a numeric case score. Optimal thresholds were determined across all devices and for each device separately using the Youden-Index. Screening performance was assessed using sensitivity, specificity, balanced accuracy, and cancer detection rate (CDR). Standard double-reading had a CDR of 5.07 (95% CI: 4.19-6.09). Using AI with a general threshold achieved a CDR of 4.81 (95% CI: 3.95-5.80), which was enhanced using device-specific thresholds to 4.90 (95% CI: 4.03-5.89). Workload for radiologists in consensus conferences was lowered by nearly one-third when using device-specific thresholds compared to a general threshold (p < 0.001). Nonetheless, AI screening performance for one of the mammography devices was poor even with device-specific thresholds. AI-based performance varied considerably among the three mammography devices, emphasizing the need for device-specific thresholds or even the withdrawal of AI use on certain devices. Awareness of different case distribution among radiologists should be raised to mitigate the risk of interpretation biases and reduced diagnostic accuracy. Lastly, training datasets must be representative not only of the target population but also of the different screening devices used in practice. Question AI performance may vary across mammography devices, raising uncertainty about device-independent thresholds. Findings AI scores differed considerably across devices, indicating the need for device-specific thresholds, but performance on one device remained poor. Critical relevance The performance of an AI algorithm in mammography screening depends on adequately defined decision thresholds, and radiologists must be aware that AI scores can vary across mammography devices to avoid interpretation biases and to maintain diagnostic accuracy.

Topics

Journal Article

Ready to Sharpen Your Edge?

Subscribe to join 11k+ peers who rely on RadAI Slice. Get the essential weekly briefing that empowers you to navigate the future of radiology.

We respect your privacy. Unsubscribe at any time.