Back to all papers

When multi-institutional dataset aggregation masks shortcut-like behavior in Foundation Models for medical imaging.

August 22, 2026pubmed logopapers

Authors

Germani E,Zeineddine F,Mourad C,Albarqouni S

Affiliations (3)

  • University of Bonn, University Hospital Bonn, Clinic for Diagnostic and Interventional Radiology, Bonn, Germany.
  • Lebanese Hospital Geitaoui, Department of Diagnostic Imaging and Interventional Therapeutics, Beirut, Lebanon.
  • University of Bonn, University Hospital Bonn, Clinic for Diagnostic and Interventional Radiology, Bonn, Germany. Electronic address: [email protected].

Abstract

Foundation Models are increasingly used in medical imaging, yet their behavior under severe domain heterogeneity remains poorly understood. We study pretrained image encoders as feature extractors for mammography classification across ten datasets spanning diverse acquisition pipelines, populations, and geographic regions. Classifiers trained on aggregated heterogeneous data exhibit dataset-dependent failure modes, with substantial disparities in error rates, confidence, and calibration across datasets. At the representation level, dataset identity is readily recoverable from Foundation Model embeddings, whereas clinical labels are less separable, indicating a low task-to-dataset signal ratio. These findings are consistent with shortcut-like behavior, where dataset-associated structure is more accessible to downstream classifiers than transferable clinical signal. Sequential vector decomposition further shows that much of the linearly recoverable task signal lies in dataset-aligned subspaces, limiting the task information that remains after suppressing dataset-related directions. An exploratory metadata-guided analysis identifies age-associated, label-dependent shifts in predicted probabilities, suggesting cohort-level candidate mediators of dataset-dependent behavior. Dataset-adversarial training, worst-group optimization, and balanced fine-tuning reduce some dataset-level disparities, but robustness gains often come at the expense of aggregate performance. Together, our results show that classifiers trained on Foundation Model features from aggregated datasets can exhibit shortcut-like behavior, highlighting the need for evaluation protocols and representation-learning objectives that explicitly assess and reduce entanglement between spurious and clinical signals.

Topics

Journal Article

Ready to Sharpen Your Edge?

Subscribe to join 11k+ peers who rely on RadAI Slice. Get the essential weekly briefing that empowers you to navigate the future of radiology.

We respect your privacy. Unsubscribe at any time.