Back to all papers

Radiologically Relevant Clinical History Summarization with Large Language Models: A Multireader Performance Study.

August 4, 2026pubmed logopapers

Authors

Serapio A,Chen TL,Tangsombatvisit B,Fields BKK,Yu Y,Guo Y,Kim SK,Miao BY,Sushil M,Hess CP,Majumdar S,Sohn JH

Affiliations (7)

  • Department of Radiology and Biomedical Imaging, University of California, San Francisco, 505 Parnassus Ave, San Francisco, CA 94143.
  • UC Berkeley-UCSF Graduate Program in Bioengineering, Berkeley, Calif.
  • School of Information Sciences, University of Illinois, Urbana-Champaign, Urbana, Ill.
  • Department of Artificial Intelligence, Ewha Womens University, Seoul, South Korea.
  • Bakar Computational Health Sciences Institute, University of California, San Francisco, Calif.
  • Department of Neurological Surgery, University of California, San Francisco, San Francisco, Calif.
  • Division of Clinical Informatics and Digital Transformation (DoC-IT), Department of Medicine, University of California, San Francisco, San Francisco, Calif.

Abstract

Background Clinical histories accompanying imaging orders guide protocol selection and diagnostic focus. However, they are often incomplete, potentially compromising diagnostic accuracy and workflow efficiency. Purpose To evaluate whether large language models (LLMs) can improve the clinical utility of provided imaging indications by leveraging clinical notes. Materials and Methods This retrospective study curated a dataset from deidentified electronic health records at the University of California San Francisco (January 2012 to August 2024), consisting of radiology reports with paired referring clinician-provided and radiologist-curated indications linked to clinical notes. The dataset was stratified across five body systems and five pathophysiologic categories to derive LLM selection and reader study internal test sets. For the reader study, 20 radiologists with 2-25 years of experience compared indications from the referring clinician, radiologist, and best-performing LLMs. Readers scored comprehensiveness, factuality, and conciseness and ranked indications for usefulness in protocoling, usefulness in interpretation, and overall ranking. Models and clinicians were compared using cumulative link mixed models with Tukey-adjusted post hoc comparisons. Results From 28 313 patients (mean age, 59 years ± 20.6 [SD]; 14 912 women), 250 examinations from 247 patients were sampled for the reader study. After nine exclusions, 241 examinations were analyzed, yielding 482 reader-examination evaluations. Indications from the best-performing proprietary (Claude 3.5 Sonnet; Anthropic) and open-source (Qwen 2.5-7B Instruct; Alibaba) LLM were rated as more comprehensive (Likert rating of 5: 37.14% and 28.42%, respectively; both <i>P</i> < .001) and factual (68.05% and 59.75%; both <i>P</i> < .001) than referring clinician indications. The proprietary LLM ranked most useful in protocoling (rank 1: 40.87%; all <i>P</i> < .001), useful in interpretation (44.61%; all <i>P</i> < .001), and overall ranking (44.19%, all <i>P</i> < .001). Comprehensiveness (65.77% of ratings; both <i>P</i> < .001) most strongly influenced overall rankings. Conclusion LLMs generated radiology-relevant indications from clinical notes that were more comprehensive and factual than clinician indications, and when generated by the proprietary LLM, were ranked most useful in protocoling and imaging interpretation. © RSNA, 2026 <i>Supplemental material is available for this article.</i> See also the editorial by Yilmaz and Cardoza-Ochoa in this issue.

Topics

Large Language ModelsElectronic Health RecordsMedical History TakingJournal Article

Ready to Sharpen Your Edge?

Subscribe to join 11k+ peers who rely on RadAI Slice. Get the essential weekly briefing that empowers you to navigate the future of radiology.

We respect your privacy. Unsubscribe at any time.