From Shadows to Signals: Exploring the Strengths and Limitations of AI in Breast Cancer Lesion Detection on Screening Mammograms Compared to Radiologists and Radiology Trainees.
Authors
Affiliations (2)
Affiliations (2)
- Discipline of Medical Imaging Sciences, Sydney School of Health Sciences, Faculty of Medicine and Health, The University of Sydney, New South Wales, Australia 2006.
- School of Health Sciences, Western Sydney University, Campbelltown, Australia 2560.
Abstract
Breast cancer screening with mammography reduces mortality, but interpretation errors persist. While artificial intelligence (AI) shows promise, most studies focus on case-level accuracy and overlook lesion-level performance. This study evaluated the AI (GMIC+CLAHE) model against radiologists and radiology trainees, considering breast density and lesion characteristics. Nine BREAST test sets of screening digital mammograms (540 cases:179 cancer, 361 normal) were analysed. AI performance was compared with 17 radiologists and 26 trainees. Case- and lesion-level performance were assessed across breast density, lesion type, and size. AI malignancy scores were dichotomised using Youden's index. Sensitivity, specificity, ROC AUC, lesion localisation, and odds ratios (ORs) were calculated. AI achieved higher case-level AUC (0.974) than radiologists (0.82±0.06;P=0.01) and trainees (0.74±0.05;P<0.0001) across breast densities. At lesion level, AI matched radiologists and outperformed trainees (OR = 1.6;P=0.03), with lesion sensitivities of 66.3% (AI), 67.0% (radiologists), and 55.0% (trainees). AI also outperformed trainees in detecting mixed-type lesions (75%-vs-55%;OR=2.4) and small lesions ≤15 mm (65.7%-vs-55%;OR=1.6). Architectural distortion had the highest miss rate (43.8%) whereas mixed-type lesions demonstrated the lowest miss rate (25.0%) by AI. AI demonstrated superior case-level accuracy compared with radiologists and trainees, but limitations in localising specific lesion types, supporting its role as an adjunct to human readers in screening and training. This study integrates case-and-lesion-level evaluation in different case characteristics, identifying key strengths and limitations for clinical implementation. Despite superior case-level accuracy, variability in lesion localisation across cancer types and sizes highlights important considerations for its integration into screening services.