Back to all papers

Systematic evaluation of foundation models for organ-level classification on CT scans.

August 27, 2026pubmed logopapers

Authors

Tagscherer J,de Boer S,van der Graaf F,Philipp L,Jacobs C,Smit EJ,Hering A

Affiliations (2)

  • Department of Medical Imaging, Radboudumc, Nijmegen, The Netherlands. [email protected].
  • Department of Medical Imaging, Radboudumc, Nijmegen, The Netherlands.

Abstract

Foundation models are increasingly used as frozen feature extractors for CT classification tasks, yet the determinants of their downstream performance remain unclear. We assess how foundation model choice, feature aggregation strategy, and abnormality type are associated with performance in organ-level abnormality classification. We further evaluate whether more expressive aggregation strategies outperform simple pooling, and whether performance varies by abnormality type. We evaluate five state-of-the-art medical imaging foundation models on binary organ-level abnormality classification (normal vs. abnormal) across six abdominal organs. Local patch-level embeddings are aggregated using six strategies, including simple pooling methods (e.g., mean pooling) and attention-based multiple instance learning. A linear classifier is trained on feature embeddings, and performance is evaluated on an annotated test set of 200 CT scans. Among the evaluated models, 3D CT-native models generally outperformed 2D multi-modal models. None of the tested aggregation strategies significantly outperformed mean pooling, including attention-based multiple instance learning (best-performing aggregation vs. mean: <math xmlns="http://www.w3.org/1998/Math/MathML"><mi>Δ</mi></math> AUC = 0.008, 95% CI <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mo>[</mo> <mo>-</mo> <mn>0.002</mn> <mo>,</mo> <mn>0.022</mn> <mo>]</mo></mrow> </math> ). In our organ-level classification setting, performance differed between abnormality type for all well-performing models (SPECTRE, TAP-CT, CT-FM), with lower AUCs observed for focal abnormalities compared to diffuse abnormalities (largest difference: <math xmlns="http://www.w3.org/1998/Math/MathML"><mi>Δ</mi></math> AUC = 0.108, 95% CI [0.080, 0.135]). This gap was not reduced by the evaluated aggregation strategies. In our experiments, downstream performance varied more across foundation models than across aggregation strategies. The choice of aggregation strategy, including attention-based multiple instance learning, did not significantly impact performance in this setting. The persistent gap for focal abnormalities suggests that current representations may insufficiently encode localized disease patterns, motivating the development of localization-aware pre-training approaches.

Topics

Journal Article

Ready to Sharpen Your Edge?

Subscribe to join 11k+ peers who rely on RadAI Slice. Get the essential weekly briefing that empowers you to navigate the future of radiology.

We respect your privacy. Unsubscribe at any time.