LLMs can generate more comprehensive and factual clinical indications for radiologists than referring clinicians.
Key Details
- 1Study evaluated LLM performance using deidentified records of 28,313 UCSF patients and 77,626 imaging exams.
- 2Both open-source (Qwen 2.5-7B Instruct) and proprietary (Claude 3.5 Sonnet) LLMs were benchmarked.
- 320 radiologists rated LLM vs. clinician-generated indications on comprehensiveness, factuality, and usefulness.
- 4Claude 3.5 Sonnet scored highest for comprehensiveness (37.14%) and factuality (68.05%), outperforming clinicians (6.64% and 50%).
- 5Comprehensiveness was the key driver for radiologists’ rankings; LLMs consistently outranked referring clinicians.
- 6Editorial stresses need for transparency, monitoring, and human oversight in deploying LLMs in radiology.
Why It Matters

Source
AuntMinnie
Related News

Rad Partners Wins $1.29M FDA Grant for AI Radiology Report Evaluation Study
Cognita Imaging, part of Rad Partners, has received a $1.29M FDA grant to develop a new framework using large language models (LLMs) to evaluate AI-generated radiology reports.

Real-World Study: Radiology AI Best in Emergency and Inpatient Settings
A commercial AI tool for intracranial aneurysm detection outperformed in inpatient and emergency settings but yielded limited benefits for outpatients in a major health system study.

New Rubric Enhances Safety of AI-Generated Radiology Summaries
Researchers developed a five-factor rubric to assess the safety and quality of AI-generated, patient-friendly radiology report summaries.