Leveraging Large Language Models to Enhance Radiology Report Readability: A Systematic Review.
Authors
Affiliations (4)
Affiliations (4)
- Department of Radiology, University of California, Irvine Building 1, 101 The City Dr S, Orange, CA, US 92868 714-456-7890.
- Department of Radiology, University of California, Irvine Building 1, 101 The City Dr S, Orange, CA, US 92868 714-456-7890. Electronic address: [email protected].
- Paul Merage School of Business, University of California, Irvine 4293 Pereira Dr, Irvine, CA US 92697 949-824-4909.
- Department of Radiology and Imaging Sciences, Emory University 1364 Clifton Rd NE BG20, Atlanta, GA, US 30322-1007 404-712-4519.
Abstract
Patients increasingly have direct access to their medical record. Radiology reports are complex and difficult for patients to understand and contextualize. One solution is to use large language models (LLMs) to translate reports into patient-accessible language. Objective This review summarizes the existing literature on using LLMs for the simplification of patient radiology reports. We also propose guidelines for best practices in future studies. A systematic review was performed following PRISMA guidelines. Studies published and indexed using PubMed, Scopus, and Google Scholar up to February 2025 were included. Inclusion criteria comprised of studies that used large language models for simplification of diagnostic or interventional radiology reports for patients and evaluated readability. Exclusion criteria included non-English manuscripts, abstracts, conference presentations, review articles, retracted articles, and studies that did not focus on report simplification. The Mixed Methods Appraisal tool (MMAT) 2018 was used for bias assessment. Given the diversity of results, studies were categorized based on reporting methods, and qualitative and quantitative findings were presented to summarize key insights. A total of 2126 citations were identified and 17 were included in the qualitative analysis. 71% of studies utilized a single LLM, while 29% of studies utilized multiple LLMs. The most prevalent LLMs included ChatGPT, Google Bard/Gemini, Bing Chat, Claude, and Microsoft Copilot. All studies that assessed quantitative readability metrics (n=12) reported improvements. Assessment of simplified reports via qualitative methods demonstrated varied results with physician vs non-physician raters. LLMs demonstrate the potential to enhance the accessibility of radiology reports for patients, but the literature is limited by heterogeneity of inputs, models, and evaluation metrics across existing studies. We propose a set of best practice guidelines to standardize future LLM research.