Back to all papers

Med-REDUCE: Representation transfer and efficiency under resolution constraints.

September 23, 2026pubmed logopapers

Authors

Bikia V,Xu S,Skrika A,Park R,Fanous A,Daneshjou R

Affiliations (3)

  • Stanford University, Department of Biomedical Data Science, Stanford, CA, USA; Stanford University, Stanford Institute for Human-Centered AI, Stanford, CA, USA. Electronic address: [email protected].
  • Stanford University, Department of Biomedical Data Science, Stanford, CA, USA.
  • Stanford University, Department of Biomedical Data Science, Stanford, CA, USA; Stanford University, Stanford Institute for Human-Centered AI, Stanford, CA, USA; Stanford University, School of Medicine, Department of Dermatology, Stanford, CA, USA.

Abstract

Vision foundation models achieve strong transfer performance but are typically evaluated at fixed spatial resolutions, leaving unclear how their representations degrade when input resolution is reduced, a common constraint in mobile clinical applications, telemedicine, and point-of-care devices. We present Med-REDUCE, a controlled framework that disentangles two axes of medical representation learning across dermatology, histopathology, and radiology: the resolution robustness of frozen foundation-model features and the parameter efficiency gained by distilling large teachers into compact students. Using linear probing across resolution ladders from 512 to 64 pixels (an 8× reduction in spatial scale), we evaluate three frozen teachers: a self-supervised DINOv3 (ViT-S/16, 21M parameters), a supervised ViT-B/16 (86M), and the medical vision-language encoder BiomedCLIP (86M), together with two distilled students (ResNet-50 and TinyViT-21M). All three teachers degrade gracefully, losing at most 0.043 domain-level AUROC from 512 to 64 pixels (0.007-0.043 across teachers and domains), although which domain degrades most steeply depends on the teacher. The smallest teacher, DINOv3, consistently achieves the strongest performance across the evaluated domains, showing that greater parameter count alone does not predict stronger frozen transfer. Under embedding distillation, the compact students match or exceed their teacher at full resolution in 39 of the 42 teacher-task-student combinations, including the multi-label radiology setting, with every shortfall at most 0.004 AUROC and within seed variability; the largest gains come from distilling each domain's weakest teacher, with BiomedCLIP students improving on the BiomedCLIP baseline by up to 0.062 AUROC. Because the 21-25M students match or beat the 86M ViT-B and BiomedCLIP teachers, distillation yields up to a ∼4× parameter reduction at no cost in accuracy, and the accuracy-parameter Pareto frontier is occupied entirely by encoders at ≤25M parameters. We release training code and evaluation pipelines to support reproducible research on efficient medical representation learning under resolution constraints.

Topics

Journal Article

Ready to Sharpen Your Edge?

Subscribe to join 11k+ peers who rely on RadAISlice. Get the essential weekly briefing that empowers you to navigate the future of radiology.

We respect your privacy. Unsubscribe at any time.