From zero-shot to fine-tuning: optimize large language models for error detection of ultrasound reports.
Authors
Affiliations (4)
Affiliations (4)
- Cancer Center, Department of Ultrasound Medicine, Zhejiang Provincial People's Hospital, Affiliated People's Hospital, Hangzhou Medical College, Hangzhou, China.
- Department of Ultrasound Medicine, Hangzhou Women's Hospital (Hangzhou Maternity and Child Health Care Hospital), Hangzhou, China.
- Department of Diagnostic Ultrasound Imaging & Interventional Therapy, Zhejiang Cancer Hospital, Hangzhou, China. [email protected].
- Cancer Center, Department of Ultrasound Medicine, Zhejiang Provincial People's Hospital, Affiliated People's Hospital, Hangzhou Medical College, Hangzhou, China. [email protected].
Abstract
High workload and inconsistent quality of ultrasound report writing may lead to diagnostic errors. This study aims to ascertain whether fine-tuned open-source large language models (LLMs) can achieve promising performance for automated quality control of Chinese ultrasound reports, when compared to proprietary LLMs. This retrospective, multi-center study included a multi-subspecialty dataset of 1800 Chinese ultrasound reports, comprising 1500 quality-controlled reports injected artificially with six predefined error types and 300 reports with naturally occurring errors. Nine proprietary LLMs (under zero-shot and few-shot paradigms) and seven open-source LLMs (under fine-tuning) were evaluated, with performance compared against that of radiologists of varying seniority. Performance was measured by detection accuracy, Macro-F1 score, precision, recall, and mean absolute error across six categories. Fine-tuned open-source LLMs, notably Qwen3-14B, achieved a detection accuracy of 0.931 and a Macro-F1 of 0.739, approaching the performance of senior radiologists. Some fine-tuned open-source LLMs maintained performance despite smaller parameter sizes and outperformed most proprietary LLMs with vastly larger parameter counts. The fine-tuned Qwen3-14B demonstrated superior recognition capability for semantic errors such as redundancy, spelling, orientation, and unit or value errors. This study demonstrates that task-specific fine-tuning enables open-source LLMs to rival proprietary LLMs and expert radiologists in Chinese ultrasound report error detection, offering a locally deployable and privacy-compliant alternative for AI-assisted clinical quality control workflows. Question Can task-specific fine-tuning improve error detection by open-source LLMs in Chinese ultrasound reports and provide an effective approach to automated report quality control? Findings Task-specific fine-tuning enabled open-source LLMs to achieve Macro-F1 scores up to 0.739, approaching that of experienced radiologists (0.764) and outperforming most proprietary LLMs. Critical relevance statement Task-specific fine-tuning allows open-source LLMs to become feasible assistants in ultrasound report quality control workflows, offering a locally deployable, privacy-preserving, and regulation-compliant solution for enhancing reporting consistency and reducing diagnostic errors in high-volume ultrasound examination procedures.