Back to all papers

VL-ThyNet: An interpretable analytical framework integrating visual foundation models and vision-language reasoning for malignancy risk assessment of C-TIRADS 4 thyroid nodules.

August 31, 2026pubmed logopapers

Authors

Chen Y,Chen X,Hu J,Tan M,Wang Y,Nie L,Wang T

Affiliations (5)

  • School of Biological Science and Medical Engineering, Hunan University of Technology, Zhuzhou, 412007, China; State Key Laboratory of Chemo and Biosensing, College of Chemistry and Chemical Engineering, Hunan University, Changsha, 410082, China.
  • School of Biological Science and Medical Engineering, Hunan University of Technology, Zhuzhou, 412007, China.
  • Department of Teaching Affairs, Zhuzhou Hospital Affiliated to Xiangya School of Medicine, Central South University, No. 116, Changjiang South Road, Tianyuan District, Zhuzhou City, Hunan Province, 412000, China.
  • Hunan Provincial Key Laboratory of the Research and Development of Novel Pharmaceutical Preparations, Changsha Medical University, Changsha, 410219, China.
  • State Key Laboratory of Chemo and Biosensing, College of Chemistry and Chemical Engineering, Hunan University, Changsha, 410082, China. Electronic address: [email protected].

Abstract

Accurate differentiation between benign and malignant C-TIRADS 4 thyroid nodules is particularly challenging because of their intermediate malignancy risk and overlapping sonographic features, yet it is essential for clinical decision-making. Herein, we propose VL-ThyNet, an interpretable analytical framework for malignancy risk assessment. VL-ThyNet integrates a DINOv2 vision foundation model with a ResNet adapter to jointly capture global semantic representations and local spatial-textural features from transverse and longitudinal views. In parallel, a vision-language model generated structured radiological descriptors together with confidence scores, providing clinically interpretable radiological information for subsequent feature fusion. Ablation experiments evaluated the effects of dual-view input, the DINOv2 encoder, and VLM-derived radiological features, showing that their contributions varied across model configurations. On the test set, VL-ThyNet achieved an accuracy of 84.8% and an AUC of 0.902, showing higher ACC and AUC than the evaluated conventional CNN- and Transformer-based models. By integrating Grad-CAM visual localization with structured VLM-based interpretation, VL-ThyNet provided clinically aligned diagnostic evidence. In a reader study, VL-ThyNet assistance was associated with improved diagnostic accuracy among radiologists and an 18.4% reduction in the average image interpretation time of the six radiologists, indicating its potential as an interpretable and efficient auxiliary analytical tool for thyroid nodule assessment.

Topics

Journal Article

Ready to Sharpen Your Edge?

Subscribe to join 11k+ peers who rely on RadAI Slice. Get the essential weekly briefing that empowers you to navigate the future of radiology.

We respect your privacy. Unsubscribe at any time.