ProtoSurv: prototype-guided adaptation of computed tomography foundation model for lung-cancer prognosis prediction.
Authors
Affiliations (5)
Affiliations (5)
- Beijing Advanced Innovation Center for Big Data-Based Precision Medicine, School of Engineering Medicine, Beihang University, Beijing, 100191, China.
- CAS Key Laboratory of Molecular Imaging, Beijing Key Laboratory of Molecular Imaging, Institute of Automation, Chinese Academy of Sciences, Beijing, 100190, China.
- Department of Radiology, First Affiliated Hospital of Jilin University, Changchun, Sichuang, 130021, China.
- Department of Radiology, First Affiliated Hospital of Jilin University, Changchun, Sichuang, 130021, China. [email protected].
- Beijing Advanced Innovation Center for Big Data-Based Precision Medicine, School of Engineering Medicine, Beihang University, Beijing, 100191, China. [email protected].
Abstract
Prognosis prediction is important for precision treatment in lung cancer. Chest computed tomography (CT) and deep learning are feasible approaches for learning imaging-based prognostic representations. However, prognostic labels are difficult to obtain because survival endpoints require long-term follow-up; thus, labeled data for supervised training is limited, increasing the risk of overfitting. Vision-language pretraining offers an approach to mitigate this problem by leveraging large-scale unlabeled CT data; however, most existing methods are not specifically designed to learn prognosis-related features. We propose ProtoSurv, a prototype-guided learning framework for adapting large-scale pretrained CT representations for lung-cancer prognosis prediction under limited supervision. ProtoSurv first pretrains a CT-report vision-language model using 409,261 chest CT sequences from 104,783 patients. It then uses limited labeled prognosis data to construct class prototypes in the pretrained feature space. These prototypes guide pseudolabel assignment, confidence-based filtering, and feature calibration, allowing the selection and refinement of task-relevant unlabeled samples before downstream prognosis modeling. We evaluated ProtoSurv based on two lung-cancer prognosis tasks: progression-free survival prediction in 507 patients receiving targeted therapy and overall survival prediction in 420 patients receiving (chemo-)radiotherapy. ProtoSurv achieved area under the curve/concordance index values of 0.765/0.684 and 0.822/0.727 for the first and second datasets, respectively, outperforming the conventional clinical models and representative deep learning baselines. These results suggest that the prototype-guided adaptation can improve the use of large-scale unlabeled CT data for prognosis modeling when labeled survival data are limited.