Tomosynthesis and Synthesized-2D Breast AI Outputs in a Biopsy-Referred Cohort: A Diagnostic Study of Agreement and Incremental Decision-Support Value.
Authors
Affiliations (4)
Affiliations (4)
- Department of Radiology, İzmir University of Economics Medical Point Hospital, İzmir, Türkiye. [email protected].
- Department of Radiology, İzmir University of Economics Medical Point Hospital, İzmir, Türkiye.
- Department of General Surgery, Izmir University of Economics Medical Point Hospital, İzmir, Türkiye.
- Department of Public Health, Faculty of Medicine, Ege University, İzmir, Türkiye.
Abstract
To determine whether two malignancy scores co-reported by a deployed breast artificial intelligence (AI) platform, one tomosynthesis-based and one synthesized-image-based, are interchangeable, and whether either adds information beyond clinical-radiological assessment. This single-center retrospective study analyzed 139 breast-level observations (84 malignant, 55 benign) from 134 biopsy-referred women, with histopathology as reference. Agreement between the two scores was assessed by paired comparison, by whether each output marked a lesion, and by threshold classification. A baseline logistic model (age, radiological suspicion, lesion size) was compared with models adding each score. Cancers negative by the tomosynthesis output at the cohort-derived threshold underwent targeted, AI-masked re-review. The outputs differed within individual breasts despite similar overall discrimination (area under the receiver operating characteristic curve [AUC] 0.728 versus 0.717; p = 0.786). Agreement was moderate (kappa = 0.53); among malignant breasts, 21 were flagged by the tomosynthesis output alone versus 6 by the synthesized output (p = 0.006). Adding either score improved model fit (p < 0.001), but discrimination gains were small (AUC 0.809 to 0.841-0.848) and decision-curve improvements uncertain. Of 30 cancers negative by the tomosynthesis output, 25 had no mammographic correlate on re-review, and breast density did not fully explain this pattern. Within this biopsy-referred cohort, the two AI outputs were not interchangeable. This supports treating co-reported outputs as distinct, context-dependent signals and not using a low AI score by itself to downgrade a suspicious finding. The reporting and governance implications require prospective multireader, multicenter validation.