Back to all papers

The Performance of ChatGPT-4o and DeepSeek-R1 in Interpreting Thyroid Nodule Ultrasound Text Reports: Multicenter Study.

July 28, 2026pubmed logopapers

Authors

Xie Y,Liu J,Zhan B,Zhang K,Li Y,Ning C

Affiliations (3)

  • Department of Ultrasound, Affiliated Hospital of Qingdao University, No. 16 Jiangsu Road, Qingdao, Shandong, 266000, China, 86 18661806751.
  • Department of Ultrasound, JiaoZhou Central Hospital of Qingdao, Qingdao, China.
  • Department of Ultrasound, Tai'an City Central Hospital, Tai'an, China.

Abstract

Although thyroid nodules are detected in up to 60% of adults on ultrasound, the vast majority are benign, creating a substantial decision-making burden compounded by heterogeneous practice guidelines. Large language models (LLMs) show promise in processing unstructured medical text and are emerging as tools for report interpretation among both clinicians and patients. However, their reliability across distinct clinical tasks in thyroid ultrasound interpretation remains poorly characterized. This study evaluates 2 LLMs, ChatGPT-4o and DeepSeek-R1, in interpreting thyroid nodule ultrasound text reports across three clinical tasks-benign-malignant differentiation, Chinese Thyroid Imaging Reporting and Data System (C-TIRADS) classification, and management recommendation-with concurrent assessment of output stability for each task. We retrospectively analyzed 1063 ultrasound text reports from 3 medical centers, including 306 with histopathological confirmation. Each nodule report was submitted to both LLMs via their consumer web interfaces using task-specific prompts, with 5 repetitions per model; final outputs were determined by mode voting. Diagnostic performance was assessed by receiver operating characteristic analysis with DeLong testing; agreement was quantified using squared weighted κ and Cohen κ; and stability was measured using Krippendorff α and Fleiss κ. For benign-malignant differentiation, DeepSeek-R1 showed higher sensitivity (0.879 vs 0.692; P<.001) and accuracy (0.729 vs 0.644; P=.008) than ChatGPT-4o. With access to images and clinical context unavailable to the LLMs, senior radiologists showed higher performance (area under the curve=0.865; accuracy=0.804). For C-TIRADS classification, DeepSeek-R1 showed substantial agreement with radiologists, exceeding ChatGPT-4o (κ=0.770 vs 0.688; Δκ=0.082, 95% CI 0.048-0.122). Both models yielded moderate, comparable agreement with clinicians on management recommendations (κ=0.606 vs 0.608). Stability was near perfect for C-TIRADS classification (α=0.864 vs 0.866) and management recommendations (κ=0.853 vs. 0.849) in both models; however, DeepSeek-R1 showed markedly greater stability than ChatGPT-4o in benign-malignant differentiation (κ=0.869 vs 0.609; Δκ=0.260, 95% CI 0.191-0.321). Both LLMs demonstrate clinical potential for thyroid nodule ultrasound report interpretation, with DeepSeek-R1 showing advantages in diagnostic accuracy, classification consistency, and output stability. However, both LLMs remained inferior to senior radiologists, suggesting their role as decision-support tools rather than stand-alone diagnostic systems. These findings provide preliminary evidence to inform the responsible integration of LLMs into thyroid imaging workflows while highlighting the need for further evaluation before patient-facing deployment.

Topics

Journal Article

Ready to Sharpen Your Edge?

Subscribe to join 11k+ peers who rely on RadAI Slice. Get the essential weekly briefing that empowers you to navigate the future of radiology.

We respect your privacy. Unsubscribe at any time.