Self-supervised multimodal swin UNETR for PET/CT segmentation of diffuse large B-cell lymphoma.
Authors
Affiliations (5)
Affiliations (5)
- Department of Medical Radiation Engineering, SR.C., Islamic Azad University, Tehran, Iran. Electronic address: [email protected].
- Department of Medical Radiation Engineering, SR.C., Islamic Azad University, Tehran, Iran. Electronic address: [email protected].
- Department of Medical Imaging, Division of Nuclear Medicine and Molecular Imaging, Geneva University Hospital, CH-1211, Geneva, Switzerland. Electronic address: [email protected].
- Department of Diagnostic Imaging, The Hospital for Sick Children (SickKids), University of Toronto, Toronto, Ontario, Canada. Electronic address: [email protected].
- Department of Medical Radiation Engineering, SR.C., Islamic Azad University, Tehran, Iran. Electronic address: [email protected].
Abstract
Accurate segmentation of diffuse large B-cell lymphoma (DLBCL) is critical for reliable PET/CT-based quantification and total metabolic tumor volume (TMTV) estimation. However, most deep learning approaches rely on limited annotated datasets and conventional convolutional architectures, which may restrict generalization and training stability. In this study, we propose a multimodal PET/CT Swin UNETR framework enhanced through large-scale self-supervised pretraining. A total of 717 FDG-PET/CT scans were collected from a single center, including 500 unlabeled volumes used for self-supervised pretraining and 217 expert-annotated cases for supervised learning. The labeled dataset was randomly divided into 174 training cases and an independent held-out test cohort of 43 cases. A standardized preprocessing pipeline with isotropic resampling and voxel harmonization was applied to ensure spatial consistency. The proposed SSL + PET/CT Swin UNETR achieved a Dice score of 0.723 (95% CI: 0.70-0.75), IoU of 0.648, HD of 15.0 mm, and HD95 of 8.8 mm on a held-out internal test set, evaluated supervised baseline models (p < 0.001, Wilcoxon signed-rank test). TMTV regression analysis demonstrated strong correlation with ground-truth measurements (R² = 0.929) with minimal systematic bias in Bland-Altman evaluation. These findings demonstrate the potential of multimodal self-supervised learning to improve automated DLBCL lesion segmentation in PET/CT images and support future research on robust quantitative image analysis.