Back to all papers

Improved Readability and Translational Instability in LLM-Generated Radiology Reports.

August 6, 2026pubmed logopapers

Authors

Mao Y,Wang C,Wang W,Zhang M

Affiliations (2)

  • Department of Radiology, China-Japan Union Hospital of Jilin University, Changchun 130000, China.
  • Department of Radiology, China-Japan Union Hospital of Jilin University, Changchun 130000, China. Electronic address: [email protected].

Abstract

Large language models (LLMs) show promise for converting complex radiology reports into patient-centric language, but inherent output instability may limit clinical application. To quantitatively assess the translational accuracy, error rates, and instability of various LLMs when generating patient-centric radiology reports, and evaluate demographic influences on report readability. This retrospective study evaluated 320 de-identified radiology reports processed by three LLMs using a two-stage (baseline and optimized) prompt engineering strategy. Two senior radiologists evaluated medical accuracy, completeness, and recommendation suitability. Readability was evaluated by 16 non-medical participants stratified by age and education. Professional radiological evaluation revealed that all tested models exhibited inherent instability, omitted information, and tended to generate risk-averse, generalized clinical recommendations. To address these limitations, optimized structured prompts significantly reduced model output variance and improved translational accuracy, with particularly prominent effects observed in DeepSeek-R1 and ChatGPT-4.0. Overall, large language models significantly enhanced the readability of radiology reports (P < 0.05), with DeepSeek-R1 achieving the best performance. However, patients' self-reported comprehension of the reports was affected by demographic characteristics. Large language models can effectively improve the readability of radiology reports, yet all such models inherently suffer from output instability and information omission. Optimized structured prompting can substantially reduce the variability of model outputs and improve the accuracy of medical text translation. Nevertheless, LLMs should currently be strictly confined to human-supervised auxiliary tools rather than applied as standalone clinical solutions.

Topics

Journal Article

Ready to Sharpen Your Edge?

Subscribe to join 11k+ peers who rely on RadAI Slice. Get the essential weekly briefing that empowers you to navigate the future of radiology.

We respect your privacy. Unsubscribe at any time.