Back to all papers

Impact of Clinician-Recorded ACR TI-RADS Features on ChatGPT-5.4-Based Thyroid Nodule Classification.

September 15, 2026pubmed logopapers

Authors

Chen Z,Chen F,Wang Y

Affiliations (3)

  • Department of Health Technology and Informatics, The Hong Kong Polytechnic University, Kowloon, Hong Kong.
  • Department of Ultrasound, The Fifth Affiliated Hospital of Sun Yat-sen University, Zhuhai, China. Electronic address: [email protected].
  • Ultrasound Department, EDAN Instruments, Inc., Shenzhen, China. Electronic address: [email protected].

Abstract

Multimodal large language models (LLMs) have shown potential in medical image analysis, but their performance in the direct interpretation of thyroid ultrasound images remains limited. Whether clinician-recorded structured sonographic features can improve LLM-based thyroid nodule classification is unclear. This study aimed to evaluate the effect of clinician-recorded structured ACR TI-RADS features on ChatGPT-5.4 thyroid nodule classification by comparing image-only, features-only, and combined image-and-feature input strategies. In this prospective cross-sectional study, 202 thyroid nodules from 153 patients with histopathological confirmation were included. ChatGPT-5.4 Thinking was evaluated under three input conditions: image-only analysis based on transverse and longitudinal grayscale ultrasound images, features-only analysis based on the five clinician-recorded ACR TI-RADS descriptors, and TI-RADS-informed analysis combining the same images with the structured descriptors. Agreement with pathology was assessed using Cohen's kappa, and diagnostic performance was evaluated using sensitivity, specificity, and accuracy. Pairwise comparisons were performed using generalized estimating equations with Holm correction for multiple testing. Of the 202 nodules, 66 were benign and 136 were malignant. Accuracy was significantly higher with TI-RADS-informed than with image-only input (84.2% vs 73.8%; Holm-adjusted P = 0.001). Features-only input also achieved an accuracy of 84.2%, which was significantly higher than that of image-only input (Holm-adjusted P = 0.005). Compared with the features-only strategy, the TI-RADS-informed strategy had lower sensitivity (91.9% vs 96.3%; Holm-adjusted P = 0.088) but higher specificity (68.2% vs 59.1%; Holm-adjusted P = 0.026). Cohen's κ values were 0.415, 0.606, and 0.625 for image-only, features-only, and TI-RADS-informed analyses, respectively. Clinician-recorded structured ACR TI-RADS features improved ChatGPT-5.4 thyroid nodule classification accuracy when provided alongside ultrasound images. However, the combined strategy did not achieve higher overall accuracy than features-only input.

Topics

Journal Article

Ready to Sharpen Your Edge?

Subscribe to join 11k+ peers who rely on RadAISlice. Get the essential weekly briefing that empowers you to navigate the future of radiology.

We respect your privacy. Unsubscribe at any time.