When multi-institutional dataset aggregation masks shortcut-like behavior in Foundation Models for medical imaging.
Authors
Affiliations (3)
Affiliations (3)
- University of Bonn, University Hospital Bonn, Clinic for Diagnostic and Interventional Radiology, Bonn, Germany.
- Lebanese Hospital Geitaoui, Department of Diagnostic Imaging and Interventional Therapeutics, Beirut, Lebanon.
- University of Bonn, University Hospital Bonn, Clinic for Diagnostic and Interventional Radiology, Bonn, Germany. Electronic address: [email protected].
Abstract
Foundation Models are increasingly used in medical imaging, yet their behavior under severe domain heterogeneity remains poorly understood. We study pretrained image encoders as feature extractors for mammography classification across ten datasets spanning diverse acquisition pipelines, populations, and geographic regions. Classifiers trained on aggregated heterogeneous data exhibit dataset-dependent failure modes, with substantial disparities in error rates, confidence, and calibration across datasets. At the representation level, dataset identity is readily recoverable from Foundation Model embeddings, whereas clinical labels are less separable, indicating a low task-to-dataset signal ratio. These findings are consistent with shortcut-like behavior, where dataset-associated structure is more accessible to downstream classifiers than transferable clinical signal. Sequential vector decomposition further shows that much of the linearly recoverable task signal lies in dataset-aligned subspaces, limiting the task information that remains after suppressing dataset-related directions. An exploratory metadata-guided analysis identifies age-associated, label-dependent shifts in predicted probabilities, suggesting cohort-level candidate mediators of dataset-dependent behavior. Dataset-adversarial training, worst-group optimization, and balanced fine-tuning reduce some dataset-level disparities, but robustness gains often come at the expense of aggregate performance. Together, our results show that classifiers trained on Foundation Model features from aggregated datasets can exhibit shortcut-like behavior, highlighting the need for evaluation protocols and representation-learning objectives that explicitly assess and reduce entanglement between spurious and clinical signals.