Development and external validation of a deep learning-based BI-RADS prediction model for breast MRI.
Authors
Affiliations (5)
Affiliations (5)
- Diagnostikum Linz, Saporoshjestraße 3, 4030 Linz, Austria.
- Diagnosezentrum Meidling, Meidlinger Hauptstraße 7/9, 1120 Vienna, Austria.
- Diagnostikum Graz, Weblinger Gürtel 25, 8054 Graz, Austria.
- AIgnostikum GmbH, Weblinger Gürtel 25, 8054 Graz, Austria.
- High-field MR Center HFMRC, Department of Biomedical Imaging and Image-guided Treatment, Medical University of Vienna, Lazarettgasse 14, 1090 Vienna, Austria. Electronic address: [email protected].
Abstract
Breast MRI is the most sensitive imaging modality for breast cancer detection but requires time-intensive interpretation by experienced breast radiologists. Artificial intelligence (AI) has emerged as a promising tool to support image interpretation; however, robust external validation across heterogeneous clinical settings remains limited. To develop a deep learning-based model for automated BI-RADS prediction from breast MRI and to externally evaluate its diagnostic performance on an independent dataset using a clinical reference standard for malignancy. In this retrospective multicenter study, a deep learning model was developed using 12,421 breast MRI examinations acquired between 2012 and 2024 at three institutions. The model was trained using radiologist-assigned BI-RADS categories as supervision targets and externally evaluated against a clinical reference standard for malignancy. Training data included examinations performed on 1.5-T and 3-T MRI systems using dynamic contrast-enhanced protocols. External validation was performed using an independent retrospective cohort of 158 breast MRI examinations acquired at an outpatient imaging center over a six-month period using different MRI hardware and imaging protocols. Diagnostic performance for malignancy was assessed using receiver operating characteristic (ROC) analysis and compared with expert radiologist assessment using the DeLong test. The external validation cohort consisted of 158 examinations, including 21 malignant and 137 benign cases. Radiologist assessment achieved an area under the ROC curve (AUC) of 0.913 (95 % confidence interval [CI]: 0.857-0.952), whereas the AI model achieved an AUC of 0.767 (95 % CI: 0.693-0.830), corresponding to a statistically significant difference of 0.146 (P = 0.0001). At a BI-RADS threshold greater than 3, the AI model identified 54.7 % of all benign examinations as low-risk at a negative predictive value (NPV) of 97.4 %. The deep learning model showed lower discrimination than routine radiologist assessment but maintained high sensitivity at the evaluated threshold. Whether its identification of low-risk examinations can safely improve workflow requires prospective evaluation.