Streamlined TI-RADS Reporting From Bilingual Speech Recognition to LLM-Assisted Conclusion Generation Using a Two-Stage Workflow.
Authors
Affiliations (2)
Affiliations (2)
- Department of Radiology and Center for Imaging Sciences, Samsung Medical Center, Sungkyunkwan University School of Medicine, Seoul, Republic of Korea.
- Department of Radiology and Center for Imaging Sciences, Samsung Medical Center, Sungkyunkwan University School of Medicine, Seoul, Republic of Korea. Electronic address: [email protected].
Abstract
To evaluate a two-stage workflow for large language model (LLM)-assisted Thyroid Imaging Reporting and Data System (TI-RADS) reporting using LLM-based bilingual automated speech recognition (ASR) and conclusion generation. From the Digital Database of Thyroid Images dataset, 149 cases with adequate image quality were included. A radiologist generated gold standard reports and dictated the reports, implementing a two-stage artificial intelligence workflow. In stage 1, commercial and LLM-based ASR models were compared for generating draft reports. Two radiologists evaluated transcription performance using correction time, preference, and word error rate during manual correction. In stage 2, an LLM (GPT-4o) generated TI-RADS scores from gold standard reports using three prompting strategies: minimal (basic-scoring instruction), detailed (comprehensive-guidelines), and detailed with retrieval-augmented generation (RAG with guideline-integration). The accuracy of American College of Radiology (ACR) TI-RADS points and categories, and Korean Thyroid Imaging Reporting and Data System (K-TIRADS) categories was assessed. Statistical analyses included Wilcoxon matched-pairs signed-rank tests and McNemar's test. In stage 1, the LLM-based ASR model demonstrated a significantly lower overall word error rate compared to the commercial ASR model (1.0 [IQR 0-2.0] vs 3.0 [IQR 2.0-3.0], P < 0.001). A radiologist required significantly shorter correction time with the LLM-based ASR model (6.38 [IQR 4.58-10.45] vs 11.39 [IQR 9.56-14.06] seconds, P < 0.001), and both radiologists showed a strong preference for the LLM-based ASR model (70.5% and 73.7%, P < 0.001). In stage 2, detailed prompting strategies significantly improved TI-RADS scoring accuracy. ACR TI-RADS point accuracy increased from 50.3% with minimal prompting to 78.5% with detailed prompting (P < 0.001). Korean TI-RADS category accuracy similarly improved from 58.4% to 79.9% (P < 0.001). In this feasibility study of a two-stage workflow, the LLM-based ASR model demonstrated superior efficiency compared with the commercial model, and detailed prompting strategies significantly improved TI-RADS scoring accuracy; however, deterministic post-processing remained essential, and prospective clinical validation is required before deployment.