Resolution-dependent self-supervised transfer in chest radiograph classification.
Authors
Affiliations (10)
Affiliations (10)
- Lab for AI in Medicine, RWTH Aachen University, Aachen, Germany. [email protected].
- Department of Diagnostic and Interventional Radiology, University Hospital RWTH Aachen, Aachen, Germany. [email protected].
- Department of Urology, Stanford University, Stanford, CA, USA. [email protected].
- Department of Radiology, Stanford University, Stanford, CA, USA. [email protected].
- Institute for Computational Genomics, Joint Research Center for Computational Biomedicine, University Hospital RWTH Aachen, Aachen, Germany.
- Department of Diagnostic and Interventional Radiology, University Hospital RWTH Aachen, Aachen, Germany.
- Else Kroener Fresenius Center for Digital Health, Technical University Dresden, Dresden, Germany.
- Department of Medicine I, University Hospital Dresden, Dresden, Germany.
- National Center for Tumor Diseases (NCT), University Hospital Heidelberg, Heidelberg, Germany.
- Lab for AI in Medicine, RWTH Aachen University, Aachen, Germany.
Abstract
Self-supervised learning (SSL) has improved visual representation learning, but its value in chest radiography remains uncertain. DINOv3 extends earlier SSL models through Gram-anchored self-distillation and explicit high-resolution adaptation. Whether these changes improve transfer learning for chest radiograph classification has not been established. We benchmarked DINOv3 against DINOv2 and supervised ImageNet initialization across seven chest radiograph datasets comprising 816,183 radiographs from pediatric and adult cohorts. ViT-B/16 and ConvNeXt-B were evaluated under full fine-tuning at 224 × 224 and 512 × 512 pixels, with targeted 1024 × 1024 experiments on three cohorts. Additional analyses examined parameter-efficient adaptation, synthetic label corruption, external validation, frozen 7B features, and computational efficiency. The primary outcome was the mean area under the receiver operating characteristic curve across labels. In adult cohorts, DINOv3 did not consistently outperform DINOv2 at 224 × 224 pixels, but became the strongest initialization at 512 × 512 pixels, especially with ConvNeXt-B. Gains were greatest for small focal and boundary-dependent abnormalities, whereas large-structure findings changed little. The pediatric cohort showed no significant benefit from DINOv3, higher resolution, or backbone choice. Scaling to 1024 × 1024 rarely improved performance and markedly increased computational cost. ConvNeXt-B remained superior to ViT-B/16 under both full and parameter-efficient adaptation. External validation preserved the 512 × 512 DINOv3 advantage, whereas synthetic label corruption showed that this benefit should not be interpreted simply as superior noise robustness. Frozen DINOv3-7B features underperformed relative to fully adapted 86 to 89M-parameter backbones. For adult chest radiograph classification, DINOv3 provides its most reliable benefit at 512 × 512 pixels, particularly with ConvNeXt-B. Fully adapted mid-sized models at 512 × 512 pixels provided the best performance-cost trade-off in our benchmark.