Back to all papers

CoM$^3$eT: A foundation model for medical image analysis through federated, multidimensional context integration

August 17, 2026arxiv logopreprint

Authors

J. Raphael Schäfer,Kai Geissler,Till Nicke,Chiara Tappermann,Karoline Heber,Eike Petersen,Habib Mergan,Lars Ole Schwen,Nick Weiss,Annika Gerken,Jan Hendrik Moltz,Tom Bisson,Isil Dogan O,Tim-Rasmus Kiehl,Norman Zerbe,Sefer Elezkurtaj,Robin S. Mayer,Nadine Flinner,Peter Wild,Isabel Dahm,Felix Peisen,Heinrich von Busch,Robert Grimm,Sebastian Arndt,Lisa Siegler,Matthias Stefan May,Antje Prasse,Natalia Artysh,Fabian Kiessling,Johannes Lotz

Abstract

Medical foundation models improve generalization when training AI models with limited labeled data, but remain confined to a single specialty, such as pathology or radiology, and to either sparse or dense outputs, such as classification or segmentation. Here, we present CoM$^3$eT (Co-representation Multidimensional Multitask Medical Transformer), a medical vision foundation model that unifies pathology and radiology, sparse and dense predictions, and two- and higher-dimensional inputs by modeling multidimensional context with attention. CoM$^3$eT outperformed other medical foundation models in an open competition spanning five tomographic, four whole-specimen, and three two-dimensional datasets, covering sparse and dense prediction tasks as well as report generation. When adapted across diverse clinical applications, training fewer than 2.5% of parameters achieved performance comparable to full fine-tuning, enabling research without access to high-performance GPU clusters. Applied to federated learning across hospitals, this approach achieved performance comparable to pooled-data training over internet connections and with consumer-grade hardware.

Topics

cs.CV

Ready to Sharpen Your Edge?

Subscribe to join 11k+ peers who rely on RadAI Slice. Get the essential weekly briefing that empowers you to navigate the future of radiology.

We respect your privacy. Unsubscribe at any time.