EquiFL-X: Communication-constrained federated calibration via label-aggregated histograms for multi-view screening mammography under acquisition shift.
Authors
Affiliations (1)
Affiliations (1)
- Department of Electronics and Automation, Vocational School, Istinye University, Ayazaga Mah., Istanbul, 34396, Sarıyer, Turkey. Electronic address: [email protected].
Abstract
Reliable probability estimates and robust performance across heterogeneous acquisition settings are important for artificial intelligence-assisted screening mammography. In federated learning, acquisition-driven non-identically distributed data can impair both discrimination and calibration across clients. This study aimed to develop a federated multi-view mammography framework that improves worst-client robustness and probability calibration while avoiding calibration procedures that require sharing instance-level prediction-label pairs. We proposed EquiFL-X, a federated framework combining client-local masked autoencoder pretraining, worst-client-oriented aggregation, and post-hoc temperature scaling from label-aggregated histogram statistics. Evaluation followed a public-data-based protocol with explicitly documented experimental settings using the Newfoundland and Labrador Breast Screening (NLBS) Dataset (NL-Breast-Screen), in which two native-resolution acquisition domains were treated as federated clients. Examination-level predictions were obtained through permutation-invariant pooling across available views. External testing was performed on the RSNA Screening Mammography dataset using the frozen model without fine-tuning, threshold re-selection, or recalibration. Uncertainty was quantified using 95% bootstrap confidence intervals and paired bootstrap comparisons. Compared with standard federated averaging on pooled held-out examinations from NLBS, EquiFL-X improved the area under the receiver operating characteristic curve from 0.880 to 0.915 and improved worst-client area under the receiver operating characteristic curve from 0.860 to 0.910, while reducing expected calibration error from 7.0% to 2.5%. At the predefined operating points, sensitivity reached 0.60 at 95% specificity and specificity reached 0.80 at 90% sensitivity. External testing on the RSNA Screening Mammography dataset was restricted to threshold-free metrics. EquiFL-X improved pooled discrimination, worst-client robustness, and probability calibration under acquisition-driven heterogeneity while limiting calibration-time information exchange to aggregated statistics. The framework may support more reliable federated decision support for screening mammography, although broader multi-institutional validation remains necessary.