Back to all papers

DINOv2-Pretrained Vision Transformers for Stroke and Heart Failure Detection from Chest X-rays with Limited Labeled Data.

September 9, 2026pubmed logopapers

Authors

Tsai MH,Veda J,Leu JS,Tsai CT

Affiliations (4)

  • Department of Electronic and Computer Engineering, National Taiwan University of Science and Technology, Taipei, 106, Taiwan (R.O.C.).
  • Department of Electronic and Computer Engineering, National Taiwan University of Science and Technology, Taipei, 106, Taiwan (R.O.C.). [email protected].
  • Department of Internal Medicine, National Taiwan University Hospital, Taipei, 100, Taiwan (R.O.C.).
  • Department of Geriatrics and Gerontology, National Taiwan University Hospital, Taipei, 100, Taiwan (R.O.C.).

Abstract

Cardiovascular and cerebrovascular diseases are leading causes of death worldwide, motivating scalable, noninvasive screening approaches. Chest X-rays (CXRs) are widely available, but current deep learning methods rely on large labeled datasets that are costly to obtain in clinical practice. We propose a three-stage framework that combines DINOv2 self-supervised learning with vision transformers (ViTs) to detect stroke and heart failure from CXRs under limited labeled data. A ViT-Small/14 backbone is first initialized with DINOv2 weights pretrained on LVD-142M dataset and then further adapted by self-supervised pretraining on 5368 unlabeled CXRs from the UCSD dataset, with the bottom 8 transformer blocks frozen to preserve general visual representations during domain adaptation. Model evaluation is performed via frozen-backbone linear probing on 2000 labeled images per task from the NTUH-iMD database. Ablation experiments confirm that each pretraining stage contributes measurable gains under linear probing. Benchmarked against five SSL baselines spanning general domain and medical domain pretraining under an identical evaluation protocol, our stroke model achieves 89.60 ± 1.39% accuracy and 89.55 ± 1.40% F1-score, outperforming all baselines including RAD-DINO, which uses nearly four times the parameters. Our heart failure model achieves 92.52 ± 1.57% accuracy, surpassing GLoRIA, CheXzero, and RAD-DINO. We further show that negative class composition critically affects performance and that normal CXR proportion should be interpreted in light of its effect on task difficulty. Attention visualizations confirm that both models focus on anatomically meaningful cardiac and pulmonary regions, supporting the clinical interpretability of the proposed framework.

Topics

Journal Article

Ready to Sharpen Your Edge?

Subscribe to join 11k+ peers who rely on RadAISlice. Get the essential weekly briefing that empowers you to navigate the future of radiology.

We respect your privacy. Unsubscribe at any time.