Back to all papers

Kaiser Permanente National Cross-Vendor Validation of Mammography Artificial Intelligence Computer-Aided Diagnosis Algorithms in a US-Representative Population

September 24, 2026medrxiv logopreprint

Authors

Arasu, V.,Gadgil, T.,Westley, M.,Pu, A.,Lee, C.,Smith Gueye, C.,Balkman, J.,Wisner, D.,Ivansco, L.,Kushi, L.

Affiliations (1)

  • Kaiser Permanente Division of Research

Abstract

BackgroundArtificial intelligence computer-aided diagnosis (AI CAD) algorithms for screening mammography have shown promise, but independent head-to-head comparisons of commercial algorithms on large, diverse cohorts remain limited. Methods786,124 mammography screening exam performed from January 1 2022 to December 31 2023 were identified from electronic medical records. Three FDA-cleared commercial algorithms (A, B, D) and the open-source academic model Mirai (C), were applied to the four standard screening views. Breast cancer within 12 months was ascertained through cancer registry linkage. Algorithms were compared by AUROC; sensitivity, specificity, and PPV at matched AI-positive rates of 5%, 10%, 15%, 20%, 25%; and the proportion of false negatives classified as AI-positive. ResultsOf 707,922 examinations with complete data, 4,436 were associated with cancer. Radiologists recalled 6.8% of examinations, with a sensitivity of 67.2%. Algorithm D had the highest AUROC (0.849), followed by algorithm B (0.818), Mirai (0.817), and algorithm A (0.815). Algorithm D also had the highest sensitivity at every AI-positive rate, from 55.3% (95% CI: 54.1, 56.6) at 5% to 77.9% (95% CI: 76.9, 79.0) at 25%, with CIs that did not overlap those of the other algorithms. No algorithm matched radiologist sensitivity at AI-positive rates of 10% or less. Algorithm B flagged a proportion of radiologist false negatives similar to that of algorithm D from 5% through 20% (20.8% vs 20.5% at 5%), despite its lower standalone sensitivity. ConclusionsAI CAD algorithms differed meaningfully in performance on the same large, diverse screening examination cohort, and an open-source academic model performed comparably to two of the three commercial products. Independent local validation at clinically relevant operating points should precede adoption of mammography AI, and algorithm choice and threshold should be matched to the intended workflow.

Topics

radiology and imaging

Ready to Sharpen Your Edge?

Subscribe to join 11k+ peers who rely on RadAISlice. Get the essential weekly briefing that empowers you to navigate the future of radiology.

We respect your privacy. Unsubscribe at any time.