Back to all papers

Comparative assessment of DeepSeek-R1 and GPT-4 for structured ultrasound reporting of adnexal masses.

August 11, 2026pubmed logopapers

Authors

Zhang M,Wang B,Li R,Cui X,Zhang X,Wang J,Jia L,Guo L,Kong Y

Affiliations (3)

  • Department of Medical Ultrasound, The First Affiliated Hospital of Xinxiang Medical University, Xinxiang, Henan, China.
  • Department of Respiratory Intensive Care, The First Affiliated Hospital of Xinxiang Medical University, Xinxiang, Henan, China.
  • Department of Medical Ultrasound, Tongji Hospital, Tongji Medical College and State Key Laboratory for Diagnosis and Treatment of Severe Zoonotic Infectious Diseases, Huazhong University of Science and Technology, Wuhan, Hubei, China.

Abstract

This study primarily evaluated the ability of two large language models (DeepSeek-R1 and GPT-4) to generate structured ultrasound reports from free-text adnexal mass reports. Secondarily, we assessed their accuracy in O-RADS classification and management recommendations, with an exploratory analysis of their performance in benign versus malignant discrimination. This study included 215 free-text ultrasound reports of adnexal masses (AMs) between July 2024 and March 2025. Each report was processed three times per model; majority voting was used to determine the final output. Three senior radiologists, blinded to model identity, evaluated structured reports against a predefined template, O-RADS categories, and management guidelines. Histopathology served as the reference standard for benign/malignant discrimination; while expert consensus served as the reference for the other endpoints. GPT-4 showed numerically better performance than DeepSeek-R1 in generating structured reports (99.5% vs. 96.7%, p = 0.07).DeepSeek-R1 outperformed GPT-4 in O-RADS accuracy (64.4% vs. 52.9%, p < 0.001) and appropriate management recommendations (66.4% vs. 58.0%, p = 0.007). Both models demonstrated good discrimination ability for benign versus malignant classification (AUC > 0.80 for both). Both DeepSeek-R1 and GPT-4 demonstrate potential for automating structured ultrasound reporting of AMs, with DeepSeek-R1 showing superior performance in classification and recommendations. However, their diagnostic performance in discriminating benign from malignant masses, while promising, requires further refinement before independent clinical application. These findings support their role as assistive tools, pending integration with image data and prospective validation.

Topics

Journal Article

Ready to Sharpen Your Edge?

Subscribe to join 11k+ peers who rely on RadAI Slice. Get the essential weekly briefing that empowers you to navigate the future of radiology.

We respect your privacy. Unsubscribe at any time.