Back to all papers

Comparative Performance of Multimodal Large Language Models in Grayscale Ultrasound-Based Classification of Thyroid Nodules.

August 22, 2026pubmed logopapers

Authors

Chen Z,Wang Y,Chen F

Affiliations (3)

  • Department of Health Technology and Informatics, The Hong Kong Polytechnic University, Kowloon, Hong Kong. Electronic address: [email protected].
  • Ultrasound Department, EDAN Instruments, Inc., Shenzhen, China.
  • Department of Ultrasound, The Fifth Affiliated Hospital of Sun Yat-sen University, Zhuhai, China.

Abstract

Multimodal large language models (LLMs) are increasingly being explored for medical image analysis, but their relative performance in thyroid ultrasound remains unclear. This study aimed to compare six publicly available multimodal LLMs for grayscale ultrasound-based classification of thyroid nodules. This prospective cross-sectional study included 178 patients with 239 thyroid nodules who underwent preoperative thyroid ultrasound followed by histopathological confirmation. Cropped grayscale ultrasound images of the maximal transverse and longitudinal views were analyzed by six publicly available multimodal LLMs: ChatGPT-5.4, Gemini 3.1 Pro, Claude Opus 4.6, Qwen3.6-Plus, Kimi K2.5, and ERNIE 5.0. All models were evaluated using the same image-input workflow and a standardized prompt, without fine-tuning or task-specific retraining. Agreement was assessed using Cohen's kappa, and diagnostic performance was evaluated using receiver operating characteristic (ROC) analysis. Radiologist benchmarks were included for comparison. All six LLMs significantly distinguished benign from malignant nodules (all P ≤ 0.001). Gemini 3.1 Pro achieved the best overall performance, with a kappa value of 0.580 and an area under the ROC curve (AUC) of 77.1% (95% CI, 71.5%-82.7%). ChatGPT-5.4 and Qwen3.6-Plus each yielded an AUC of 73.5%, and Kimi K2.5 achieved an AUC of 71.3%. Claude Opus 4.6 and ERNIE 5.0 showed lower overall performance, with AUCs of 65.7% and 59.9%, respectively. The senior radiologist achieved higher diagnostic performance than all six LLMs. Publicly available multimodal LLMs showed measurable but heterogeneous performance in grayscale ultrasound-based thyroid nodule classification. Gemini 3.1 Pro demonstrated the best overall results, but none of the models matched senior radiologist-level performance.

Topics

Journal Article

Ready to Sharpen Your Edge?

Subscribe to join 11k+ peers who rely on RadAI Slice. Get the essential weekly briefing that empowers you to navigate the future of radiology.

We respect your privacy. Unsubscribe at any time.