Comparative assessment of DeepSeek-R1 and GPT-4 for structured ultrasound reporting of adnexal masses.
Authors
Affiliations (3)
Affiliations (3)
- Department of Medical Ultrasound, The First Affiliated Hospital of Xinxiang Medical University, Xinxiang, Henan, China.
- Department of Respiratory Intensive Care, The First Affiliated Hospital of Xinxiang Medical University, Xinxiang, Henan, China.
- Department of Medical Ultrasound, Tongji Hospital, Tongji Medical College and State Key Laboratory for Diagnosis and Treatment of Severe Zoonotic Infectious Diseases, Huazhong University of Science and Technology, Wuhan, Hubei, China.
Abstract
This study primarily evaluated the ability of two large language models (DeepSeek-R1 and GPT-4) to generate structured ultrasound reports from free-text adnexal mass reports. Secondarily, we assessed their accuracy in O-RADS classification and management recommendations, with an exploratory analysis of their performance in benign versus malignant discrimination. This study included 215 free-text ultrasound reports of adnexal masses (AMs) between July 2024 and March 2025. Each report was processed three times per model; majority voting was used to determine the final output. Three senior radiologists, blinded to model identity, evaluated structured reports against a predefined template, O-RADS categories, and management guidelines. Histopathology served as the reference standard for benign/malignant discrimination; while expert consensus served as the reference for the other endpoints. GPT-4 showed numerically better performance than DeepSeek-R1 in generating structured reports (99.5% vs. 96.7%, p = 0.07).DeepSeek-R1 outperformed GPT-4 in O-RADS accuracy (64.4% vs. 52.9%, p < 0.001) and appropriate management recommendations (66.4% vs. 58.0%, p = 0.007). Both models demonstrated good discrimination ability for benign versus malignant classification (AUC > 0.80 for both). Both DeepSeek-R1 and GPT-4 demonstrate potential for automating structured ultrasound reporting of AMs, with DeepSeek-R1 showing superior performance in classification and recommendations. However, their diagnostic performance in discriminating benign from malignant masses, while promising, requires further refinement before independent clinical application. These findings support their role as assistive tools, pending integration with image data and prospective validation.