Domain-aligned self-supervision: CT-specialized DINOv3 enhances PET/CT tumor segmentation.
Authors
Affiliations (7)
Affiliations (7)
- Department of Biomedical Engineering, Georgia Institute of Technology, 225 North Ave, Atlanta, Georgia, 30332-0002, United States.
- The University of Chicago Department of Radiation and Cellular Oncology, 5758 South Maryland Avenue, Chicago, Illinois, 60637, United States.
- Radiation oncology, Emory University School of Medicine, 1365 Clifton Rd N E, Atlanta, Georgia, 30322, United States.
- Department of Radiation Oncology, Emory University, 1365 CLIFTON RD NE ATLANTA, Atlanta, Georgia, 30322, United States.
- Department of Radiation Oncology, Emory University, 1520 Clifton Rd, Atlanta, Georgia, 30322, United States.
- Department of Radiology Oncology, Emory University, 1365-C Clifton Road NE, Atlanta, Georgia, 30322, United States.
- Department of Biomedical Engineering, National Taiwan University, 49 Fanglan Road, Taipei, 10617, Taiwan.
Abstract
This study investigates how domain-aligned self-supervised pretraining influences downstream PET/CT tumor segmentation, with a focus on disentangling the roles of domain proximity and pretraining dataset scale.

Approach: We compared nnU-Net segmentation models initialized from random weights or from DINOv3 encoders pretrained on natural images, X-ray mammography, CT, or PET, using datasets of different sizes. To disentangle the effects of pretraining domain and dataset size, CT-DINOv3 encoders were pretrained using different numbers of 2D CT slices, including subsets matched in size to the pretraining datasets from other domains. Models were evaluated on two PET/CT benchmarks, autoPET (whole-body, FDG/PSMA) and HECKTOR (head-and-neck primary tumors and metastatic lymph nodes), using five-fold cross-validation with frozen encoders.

Main results: CT-pretrained models achieved the highest Dice scores for lymph nodes, FDG lesions, and PSMA lesions. For primary tumors, PET pretraining produced the highest observed Dice score, followed by CT pretraining, suggesting the potential benefit of in-domain representations when PET data are sufficient. Notably, PET pretraining exceeded CT pretraining using equally sized datasets for PSMA but remained below models pretrained on larger CT datasets for both PSMA and FDG targets. In contrast, natural-image pretraining did not outperform from-scratch training despite its much larger pretraining scale, indicating that scale alone is insufficient when domain proximity is low. Across comparisons, strong transfer reflected the combination of domain proximity and sufficient pretraining scale, especially for CT-based models with broad data availability.

Significance: These results demonstrate that both domain proximity and dataset scale jointly determine the effectiveness of self-supervised pretraining. While higher domain alignment improves performance at comparable scale, sufficient data is required to realize this benefit. CT-based pretraining provides a favorable balance between domain relevance and data availability, leading to robust performance across tasks. This study provides practical guidance for selecting pretraining strategies in PET/CT segmentation.