Back to all papers

Comparing Radiologists' and Artificial Intelligence Performance Detecting Suspicious Microcalcifications on Screening Mammograms: A Pilot Cross-Sectional Study.

July 27, 2026pubmed logopapers

Authors

Lewis SJ,Wells JB,Jiang Z,Awwad DA,Barron ML,Trieu PDY

Affiliations (2)

  • Faculty of Medicine and Health, The University of Sydney, NSW, Australia.
  • Faculty of Health, Western Sydney University, NSW, Australia.

Abstract

IntroductionEarly breast cancer detection through periodic screening is crucial for reducing mortality due to later stage detection. Microcalcifications are common mammographic findings, present in both malignant lesions, benign pathologies, and normal tissues. This study aimed to assess radiologists' observer performance in determining suspicious calcifications requiring recall, and to compare reader performance to an in-house AI model trained and tested on Australian screening mammograms.MethodsIn this pilot proof-of-concept cross-sectional study, radiologists (n=27), breast physicians (n=2), and final year radiology trainees (n=6), completed the same mammographic test set consisting of 30 mammographic cases displaying different types of calcifications (10 breast cancer, 20 normal/benign). An in-house trained artificial intelligence model (<i>Sydney-GMIC)</i> was also applied to the same test set. Performance between readers was compared to AI via Spearman Rank-Order Correlation test. Work experience and caseload trends were compared using independent T tests and Mann-Whitney-U.ResultsSensitivity was significantly higher in radiologists with ≤ 10 years of experience compared to radiology trainees (72.2% vs 53.3%, <i>p</i>=0.042). Furthermore, readers with higher cases read per week (CPW) (i.e. 101-200 CPW compared to 0-20 CPW) had a decreased specificity (58.8% vs 74.6%, <i>p</i>=0.041), but higher sensitivity (68.8% vs 59.2%, <i>p</i>=0.12). In this pilot dataset, the <i>Sydney-GMIC</i> AI model demonstrated higher sensitivity and specificity than the radiologist means. The cases perceived as difficult by AI differed substantially from those challenging for human readers.ConclusionsThis study highlights the challenging nature of recalling calcifications from screening mammograms only, with variable performance among readers. The findings of this small pilot study are exploratory in nature, but the AI performance signals the potential utility of AI models in mammographic analysis of screening cases to progress to recall.

Topics

Journal Article

Ready to Sharpen Your Edge?

Subscribe to join 11k+ peers who rely on RadAI Slice. Get the essential weekly briefing that empowers you to navigate the future of radiology.

We respect your privacy. Unsubscribe at any time.