VL-ThyNet: An interpretable analytical framework integrating visual foundation models and vision-language reasoning for malignancy risk assessment of C-TIRADS 4 thyroid nodules.
Authors
Affiliations (5)
Affiliations (5)
- School of Biological Science and Medical Engineering, Hunan University of Technology, Zhuzhou, 412007, China; State Key Laboratory of Chemo and Biosensing, College of Chemistry and Chemical Engineering, Hunan University, Changsha, 410082, China.
- School of Biological Science and Medical Engineering, Hunan University of Technology, Zhuzhou, 412007, China.
- Department of Teaching Affairs, Zhuzhou Hospital Affiliated to Xiangya School of Medicine, Central South University, No. 116, Changjiang South Road, Tianyuan District, Zhuzhou City, Hunan Province, 412000, China.
- Hunan Provincial Key Laboratory of the Research and Development of Novel Pharmaceutical Preparations, Changsha Medical University, Changsha, 410219, China.
- State Key Laboratory of Chemo and Biosensing, College of Chemistry and Chemical Engineering, Hunan University, Changsha, 410082, China. Electronic address: [email protected].
Abstract
Accurate differentiation between benign and malignant C-TIRADS 4 thyroid nodules is particularly challenging because of their intermediate malignancy risk and overlapping sonographic features, yet it is essential for clinical decision-making. Herein, we propose VL-ThyNet, an interpretable analytical framework for malignancy risk assessment. VL-ThyNet integrates a DINOv2 vision foundation model with a ResNet adapter to jointly capture global semantic representations and local spatial-textural features from transverse and longitudinal views. In parallel, a vision-language model generated structured radiological descriptors together with confidence scores, providing clinically interpretable radiological information for subsequent feature fusion. Ablation experiments evaluated the effects of dual-view input, the DINOv2 encoder, and VLM-derived radiological features, showing that their contributions varied across model configurations. On the test set, VL-ThyNet achieved an accuracy of 84.8% and an AUC of 0.902, showing higher ACC and AUC than the evaluated conventional CNN- and Transformer-based models. By integrating Grad-CAM visual localization with structured VLM-based interpretation, VL-ThyNet provided clinically aligned diagnostic evidence. In a reader study, VL-ThyNet assistance was associated with improved diagnostic accuracy among radiologists and an 18.4% reduction in the average image interpretation time of the six radiologists, indicating its potential as an interpretable and efficient auxiliary analytical tool for thyroid nodule assessment.