Back to all papers

Evaluating large language models for lay summaries of radiology reports using tailored prompting strategies and mixed-method assessment.

September 17, 2026pubmed logopapers

Authors

Tasneem N,van der Pol CB,Zahoor A,Juggath N,McGowan K,Lokker C,Saha A

Affiliations (7)

  • MSc eHealth Program, McMaster University, Hamilton, Ontario, Canada.
  • Department of Medical Imaging, McMaster University, Hamilton, Ontario, Canada.
  • Department of Diagnostic Imaging, Juravinski Hospital and Cancer Centre, Hamilton Health Sciences, Hamilton, Ontario, Canada.
  • Escarpment Cancer Research Institute, Hamilton Health Sciences and McMaster University, Hamilton, Ontario, Canada.
  • Health Information Research Unit, Department of Health Research Methods, Evidence, and Impact, Faculty of Health Sciences, McMaster University, Hamilton, Ontario, Canada.
  • Department of Oncology, McMaster University, Juravinski Cancer Centre, Hamilton, Ontario, Canada.
  • CentRE for dAta science and digiTal hEalth (CREATE), Hamilton Health Sciences, Hamilton, Ontario, Canada.

Abstract

Radiology reports are often filled with medical jargon that limits patient understanding. Lay summaries can improve understanding but are time-consuming for healthcare providers to create. The objective of this study is to explore the use of tailored prompts for five Large Language Models (LLMs) in generating lay summaries from radiology reports. Using 100 reports from the publicly available "BioNLP 2023 report summarization" dataset, lay summaries were generated by each LLM, under select prompting styles [Few-Shot (GPT-4), Generated Knowledge (GPT-4o mini, Gemini 1.5 - Pro, Gemini 1.5 - Flash), and Zero-Shot (Llama 3.1)] informed by a pilot work. The summaries were evaluated using a mixed-method framework: subjective assessment (Likert statements) by blinded experts (n = 2 radiology fellows) and Large Reasoning Models (LRMs) [(Gemini 2.5 - Pro (LRM 1); GPT-oss-120b (LRM 2)], and readability metrics (Flesch-Kincaid Grade Level and Flesch Reading Ease). Using percentage agreement of Likert statements, the LLM-prompt combinations' performances were ranked, and Friedman and post-hoc Nemenyi tests were conducted. Gemini 1.5 - Flash and - Pro (generated knowledge) were rated highest by human experts and LRMs for generating actionable lay summaries that require minimal supervision [P < 4.97 × 10-2 (Rater 1); P < 9.03 × 10-21 (Rater 2), P < 6.90 × 10-15 (LRM 1), P < 2.760 × 10-5 (LRM 2). GPT-4 (few-shot) achieved the highest human-rated accuracy (98%), while Gemini 1.5 - Flash (LRM 1-rated: 95%) and Gemini 1.5 - Pro (LRM 2-rated: 91%) ranked first in LRM-rated accuracy. Gemini 1.5 - Pro produced the most accessible summaries (Flesch-Kincaid Grade Level: 7.55 ± 1.38, Flesch Reading Ease: 67.84 ± 7.78). Strong agreement was observed between experts and LRMs [0.96% (LRM 1) and 3.4% (LRM 2) complete disagreement]. Overall, this study highlights Gemini-models with generated knowledge prompts and the potential of LRM evaluators in assessing LLM-generated lay summaries.

Topics

Journal Article

Ready to Sharpen Your Edge?

Subscribe to join 11k+ peers who rely on RadAISlice. Get the essential weekly briefing that empowers you to navigate the future of radiology.

We respect your privacy. Unsubscribe at any time.